Heterogeneous graph clustering using inter-point mutual information criterion

By constructing a heterogeneous node network and using PMI scores for asset clustering, the problem of low accuracy in identifying malicious assets in the content distribution system was solved, achieving rapid and accurate asset labeling and reducing the risk of policy violations.

CN115280305BActive Publication Date: 2026-04-07GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-02-24
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In content distribution systems, existing technologies struggle to accurately identify assets associated with malicious or suspicious content sources, leading to frequent policy violations, increased risks, and high costs for objections or litigation.

Method used

By constructing a heterogeneous network of nodes, node clusters are identified using Point-to-Point Mutual Information (PMI) scores. High-precision asset clustering is performed based on policy labeling, automatically marking malicious or suspicious assets and reducing the frequency and risk of policy violations.

Benefits of technology

It enables rapid identification and labeling of malicious or suspicious assets, reduces the frequency of policy violations and the cost of objections or litigation, and improves the security and reliability of the content distribution system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115280305B_ABST
    Figure CN115280305B_ABST
Patent Text Reader

Abstract

A system and method are provided for implementing policies in a computing environment for content distribution using point-to-point mutual information (PMI) based clustering. The system can maintain a network of nodes representing multiple assets. When an asset is detected to be associated with a policy tag, the system can identify the asset's attributes and calculate a PMI score indicating whether nodes in the network that share the attribute belong to a single content source. When the PMI score exceeds a predefined threshold, the system can identify a cluster of nodes that include the shared attribute. The system can then label the cluster as, for example, as associating it with a content source associated with the policy tag.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] In computer networking environments such as the Internet, third-party content providers offer third-party content items for display on end-user computing devices. These third-party content items, such as text, software programs, images, and / or videos, can be displayed on web pages associated with their respective publishers, client applications, or game applications. Summary of the Invention

[0002] At least one aspect relates to a system including at least one processor and a memory storing computer-executable instructions. The computer-executable instructions, when executed by the at least one processor, enable the at least one processor to: maintain a heterogeneous network of nodes comprising multiple nodes and edges connecting corresponding node pairs. Each of the multiple nodes may represent a corresponding asset among multiple assets corresponding to multiple content sources. The multiple assets may include at least one asset of a first asset type and at least one asset of a second asset type. The at least one processor may detect that the first asset among the multiple assets has a tag associated with a policy of a content distribution system. The at least one processor may identify a first node associated with the first asset in the heterogeneous network of nodes. The at least one processor may identify combinations of two or more attributes of the first node. The at least one processor may compute a corresponding point-to-point mutual information (PMI) score for a subset of nodes in the heterogeneous network of nodes, the PMI score indicating, based on node pairs, the likelihood that a node in the node subset having the combination of said two or more attributes is associated with a single content source. The at least one processor may use the PMI score associated with the node subset to identify a cluster of nodes in the heterogeneous network of nodes including the combination of said two or more attributes. At least one processor can store the association between node clusters and tags in one or more data structures, where tags are based on a first asset having a label associated with a content distribution system's strategy. Tags can be used to classify a first set of assets among multiple assets corresponding to a node cluster.

[0003] At least one aspect relates to a method comprising a data processing system, the data processing system including one or more processors, the data processing system maintaining a heterogeneous network of nodes, the heterogeneous network of nodes including multiple nodes and edges connecting corresponding node pairs. Each of the multiple nodes may represent a corresponding asset among multiple assets corresponding to multiple content sources. The multiple assets may include at least one asset of a first asset type and at least one asset of a second asset type. The method may include at least one processor detecting that the first asset among the multiple assets has a tag associated with a policy of a content distribution system. The method may include at least one processor identifying a first node associated with the first asset in the heterogeneous network of nodes. The method may include at least one processor identifying a combination of two or more attributes of the first node. The method may include: at least one processor calculating a corresponding point-to-point mutual information (PMI) score for a subset of nodes in the heterogeneous network of nodes, the PMI score indicating, based on node pairs, the likelihood that a node in the node subset having the combination of said two or more attributes is associated with a single content source. The method may include: at least one processor using the PMI score associated with the node subset to identify a cluster of nodes in the heterogeneous network of nodes including the combination of said two or more attributes. The method may include: at least one processor storing the association between node clusters and tags in one or more data structures, the tags being based on a first asset having a tag associated with a content distribution system strategy. The tags are used to classify a first set of assets among multiple assets corresponding to the node clusters.

[0004] At least one aspect relates to a non-transitory computer-readable medium storing computer-executable instructions, which, when executed by at least one processor, cause the at least one processor to maintain a heterogeneous network of nodes, the heterogeneous network of nodes comprising a plurality of nodes and edges connecting corresponding pairs of nodes. Each of the plurality of nodes may represent a corresponding asset among a plurality of assets corresponding to a plurality of content sources. The plurality of assets may include at least one asset of a first asset type and at least one asset of a second asset type. The at least one processor may detect that the first asset among the plurality of assets has a tag associated with a policy of a content distribution system. The at least one processor may identify a first node associated with the first asset in the heterogeneous network of nodes. The at least one processor may identify a combination of two or more attributes of the first node. The at least one processor may compute a corresponding point-to-point mutual information (PMI) score for a subset of nodes in the heterogeneous network of nodes, the PMI score indicating, based on node pairs, the likelihood that a node in the subset of nodes having the combination of the two or more attributes is associated with a single content source. The at least one processor may use the PMI score associated with the subset of nodes to identify a cluster of nodes in the heterogeneous network of nodes including the combination of the two or more attributes. At least one processor can store the association between node clusters and tags in one or more data structures, whereby the tags are based on a first asset having a label associated with a content distribution system strategy. The tags can be used to classify a first set of assets among multiple assets corresponding to a node cluster.

[0005] These and other aspects and implementations are discussed in detail below. The foregoing information and the following detailed description include illustrative examples of various aspects and implementations, and provide an overview or framework for understanding the nature and characteristics of the claimed aspects and implementations. The accompanying drawings provide illustration and further understanding of the various aspects and implementations, and are incorporated in and constitute a part of this specification. It should be understood that aspects and implementations may be combined, and features described in the context of one aspect or implementation may be implemented in the context of other aspects. Attached Figure Description

[0006] The accompanying drawings are not to scale. The same reference numerals and names in the various drawings indicate the same elements. For clarity, not every component can be labeled in every drawing. In the drawings:

[0007] Figure 1 This is a block diagram illustrating one implementation of an environment for implementing strategies associated with a content distribution system, according to an illustrative implementation.

[0008] Figure 2 This is a diagram illustrating an example network of nodes according to an illustrative implementation.

[0009] Figure 3 This illustrates an illustrative embodiment. Figure 2 A graph showing the clustering results within the node network.

[0010] Figure 4 This is a flowchart illustrating a method for implementing a strategy associated with content distribution according to an illustrative embodiment.

[0011] Figure 5 The overall architecture of an illustrative computer system according to an illustrative implementation is shown. Detailed Implementation

[0012] The following is a more detailed description of various concepts and their implementations related to methods, apparatuses, and systems for implementing strategies in a computer environment for content distribution using point-to-point mutual information (PMI) based clustering methods. The various concepts introduced above and discussed in more detail below can be implemented in any of a variety of ways, as the described concepts are not limited to any particular implementation. For example, while content or content items in this document may generally be referred to as advertisements, it should be understood that content or content items can be any suitable content.

[0013] Content distribution systems that provide third-party content (e.g., text, software programs, images, and / or videos), such as advertising distribution systems, can set policies for activities performed or associated with third-party content providers. For example, content distribution policies can prohibit activities such as the distribution of malware, deceptive practices, or other unacceptable business practices. Content distribution systems can monitor activities associated with the assets of third-party content providers to determine whether the third-party content provider has engaged in prohibited activities, suspicious activities, or other activities categorized or flagged by one or more policies of the content distribution system. The assets of a third-party content provider may include, for example, third-party content provider accounts, websites, web domains, login pages, hosts, host IP addresses, resources or data files loaded from content items from hosts or login pages, or payment information. Third-party content provider accounts can include accounts with data associated with one or more activities.

[0014] A content delivery system can maintain a heterogeneous network of nodes, where each node represents an asset associated with or created by a corresponding third-party provider. This heterogeneous network can include edges, each connecting a corresponding pair of nodes. Each edge in the heterogeneous network can indicate a relationship (or correlation) between the corresponding pair of nodes or pairs of corresponding assets. For example, an edge in the node network could include an edge between a third-party content provider account and a website, indicating that at least one of the account's content items has a login page on that website. An edge in the node network could include an edge between a third-party content provider account and corresponding payment information. An edge in the node network could include an edge between a host and a corresponding IP address. An edge in the node network could include an edge between a host and resources loaded from that host, and so on. Given that the number of third-party content providers subscribing to the content delivery system and / or the number of assets associated with any given third-party content provider can be relatively large, the total number of nodes in the heterogeneous network can be very large, for example, millions.

[0015] When a content delivery system detects policy violations or policy tagging activity associated with an asset—such as a third-party content item, landing page, or resource—the system can tag the asset, stop providing it, or take other actions. In cases of serious policy violations, such as malware distribution, deception, or other unacceptable business practices, the system may suspend the entire third-party content provider account and flag or blacklist the associated website, IP address, and / or other resources. To prevent malicious third-party content providers from reusing suspicious assets and to prevent or delay the reinstatement of accounts that have violated policies from suspension, it is important for the content delivery system to identify and block all assets of the specific content source (e.g., the third-party content provider or its agents) that submitted the policy violation.

[0016] Upon detecting a policy violation or tagging activity or behavior associated with a first asset, the content distribution system can use a heterogeneous network of nodes to automatically identify, tag, and / or label all other assets related to the first asset (e.g., those sharing the same content source as the first asset). Thus, the content distribution system can automatically enforce the corresponding policy. However, policy-based automatic tagging, classification, or labeling of assets requires very high accuracy in identifying assets associated with a given content source to avoid the risk of objections or appeals. Frequent objections or appeals from content sources—e.g., third-party content providers or their agents—can be costly to handle, for example, in terms of resource allocation and may lead to reputational damage.

[0017] Another aspect of policy-based tagging, classification, or labeling of assets is the time gap between the creation or first use of a malicious, suspicious, or non-compliant (e.g., non-compliant with a given policy) asset and its tagging, labeling, or suspension. The purpose of a content delivery system or its data processing system is to minimize this time gap in order to reduce the frequency and risk or potential risk of policy violations. For example, rapid detection and suspension (or tagging) of malicious or non-compliant assets or malicious or non-compliant content sources can reduce the number of violations submitted by that asset or content source, and thus reduce the risk to client devices receiving content from the content delivery system. Even without suspension, rapid tagging of an asset or corresponding content source based on a given policy can allow, for example, further investigation or monitoring of activity associated with the asset or content source, warning the content source, or taking other preventative measures. On the other hand, waiting for a large number of assets associated with a given asset cluster to submit violations or engage in non-compliant or suspicious activities or behaviors before suspending or tagging the entire cluster can increase the rate at which client devices are exposed to such violations or suspicious activities. In some cases, by the time the cluster is suspended, substantial damage to the client device may have already occurred. This reduction in the time gap between the creation or first use of malicious or suspicious assets and their suspension also reduces resource usage, as service assets that violate regulations or do not conform to a given policy become less frequent.

[0018] This disclosure describes methods and systems for relatively high-precision clustering of assets corresponding to a single content source. The data processing system can utilize a heterogeneous network of nodes and generate asset clusters corresponding to content sources that submit detected policy violations in real time. The methods and systems described herein allow for the suspension or tagging of assets associated with a malicious, suspicious, or unqualified content source once a policy tag is detected associated with an asset from that malicious, suspicious, or unqualified content source. Specifically, upon detecting that a first asset has a policy tag (or is associated with a policy tag), the data processing system can identify a cluster of nodes in a heterogeneous network of nodes that represents assets sharing a combination of attributes with the first asset or another asset associated with the first asset. For example, the data processing system can start with a first domain associated with the first asset and identify all domains that share a combination of attributes with the first domain. The combination of attributes can include, for example, payment information, IP addresses, resources, or combinations thereof. Using more than one attribute in a combination allows for clustering using multidimensional relationships between assets.

[0019] Policy tags can include tags indicating policy violations, whether an asset contains certain information, whether any complaints about the asset have been received from users, whether the asset has been tagged by one or more other systems, asset or activity classifications, or combinations thereof. For example, policy tags can come from a finite set, such as a binary set indicating whether an asset contains certain information, or a ternary set with third tags indicating non-deterministic decisions. The cardinality of the candidate tag set can be greater than 3.

[0020] For each node within a subset of nodes in a heterogeneous network of nodes (e.g., nodes representing a domain), the data processing system can calculate a point-to-point mutual information (PMI) score indicating the likelihood of a combination of attributes being associated with that node. The data processing system can then identify nodes with corresponding PMI scores exceeding a predefined threshold as forming clusters representing assets sharing the same combination of attributes. A cluster can be viewed as a collection of assets sharing the same content source (e.g., a third-party content provider) as the first asset.

[0021] According to an example aspect of this disclosure, systems and methods for automatically enforcing content distribution policies may include a data processing system that maintains a heterogeneous network of nodes, the heterogeneous network comprising multiple nodes and edges connecting corresponding node pairs. Each of the multiple nodes may represent a corresponding asset among multiple assets corresponding to multiple content sources. The multiple assets may include at least one asset of a first asset type and at least one asset of a second asset type. The data processing system may detect that a first asset among the multiple assets has or is associated with a policy tag defined by or associated with a policy of the content distribution system. The data processing system may identify a first node associated with the first asset in the heterogeneous network of nodes. The data processing system may identify combinations of two or more attributes of the first node. The data processing system may compute a corresponding point-to-point mutual information (PMI) score for each node in a subset of nodes in the heterogeneous network of nodes, the PMI score indicating the likelihood that a combination of two or more attributes is associated with a node based on node pairs. The data processing system may use the PMI scores associated with the node subset to identify node clusters in the heterogeneous network of nodes that include combinations of said two or more attributes. The data processing system can store the association between node clusters and tags in one or more data structures. Tags are based on a first asset having a policy tag (or being associated with) the content distribution system's policy. Tags can be used to classify a first set of assets among multiple assets corresponding to a node cluster.

[0022] Figure 1This is a block diagram depicting one embodiment of an environment for maintaining the integrity of content distribution among multiple computing devices, according to an illustrative implementation. Environment 100 may include at least one data processing system 110, one or more content provider computing devices 115, one or more publisher computing devices 120, one or more client devices 125, and a network 105. The data processing system 110, one or more content provider computing devices 115, one or more publisher computing devices 120, and one or more client devices 125 may be communicatively coupled to each other via the network 105.

[0023] Data processing system 110 may include at least one processor (or processing circuitry) and memory. The memory may store computer-executable instructions that, when executed by the processor, cause the processor to perform one or more operations described herein. The processor may include a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a combination thereof. The memory may include, but is not limited to, electronic, optical, magnetic, or any other storage or transmission device capable of providing program instructions to the processor. The memory may also include floppy disks, CD-ROMs, DVDs, magnetic disks, memory chips, ASICs, FPGAs, read-only memory (ROM), random access memory (RAM), electrically erasable ROM (EEPROM), erasable programmable ROM (EPROM), flash memory, optical media, or any other suitable memory from which the processor can read instructions. Instructions may include code from any suitable computer programming language. Data processing system 110 may include one or more computing devices or servers capable of performing various functions. In some embodiments, data processing system 110 may include an advertising auction system configured to host auctions. In some embodiments, data processing system 110 may not include an advertising auction system, but may be configured to communicate with an advertising auction system via network 105.

[0024] Network 105 may include computer networks such as the Internet, Local Area Network (LAN), Wide Area Network (WAN), Metropolitan Area Network, one or more intranets, satellite networks, cellular or network networks, optical networks, other types of data networks, or combinations thereof. Data processing system 110 may communicate via network 105 with one or more content provider computing devices 115, one or more content publisher computing devices 120, or one or more client devices 125. Network 105 may include any number of network devices, such as gateways, switches, routers, modems, repeaters, and wireless access points. Network 105 may also include computing devices such as computer servers. Network 105 may also include any number of hardwired and / or wireless connections.

[0025] One or more content provider computing devices 115 may include computer servers, personal computers, handheld devices, smartphones, or other computing devices operated by content provider entities such as advertisers or their agents. One or more content provider computing devices 115 may provide content, such as text content, image content, video content, animated content, software program content, content items and / or Uniform Resource Locators (URLs), and other types of content, to data processing system 110 for display on information resources. Specifically, one or more content provider computing devices 115 for a given content provider may be a source of content items or content used to generate content items for that content provider. Content items may be used for display in information resources rendered on client device 125, such as websites, web pages of search results, client applications, game applications or platforms, open-source content sharing platforms (e.g., YouTube, DailyMotion, or Vimeo), or social media platforms.

[0026] Data processing system 110 may provide one or more user interfaces accessible via content provider computing device 115 to allow content provider entities to, for example, generate corresponding content provider accounts or generate corresponding activities. The user interfaces (or more) may allow content provider entities to upload corresponding content to data processing system 110 or other remote systems, or provide corresponding payment information. Typically, the user interfaces (or more) may allow each content provider entity to instruct on corresponding assets for distributing the entity's content to client device 125. As used herein, the assets of a content provider entity may include content provider accounts, content distribution activities, payment information, domain names, hosts, websites or web pages including login pages, content items (e.g., software programs, images, video clips, text, animation clips, etc.), or resources accessible from websites, domains, or hosts, as well as other assets of the content provider.

[0027] Content publisher computing device 120 may include a server or other computing device operated by the content publishing entity to provide primary content for display via network 105. The primary content may include websites, web pages, client applications, game content, or social media content for display on client device 125. The primary content may include search results provided by a search engine. Pages, video clips, or other units of the primary content may include executable instructions, such as instructions associated with content (or advertising) slots, that cause client device 125 to request third-party content from data processing system 110 or other remote systems when the primary content is displayed on the client device. In response to such a request, the data processing system or other remote system may run an auction to determine which content items to provide to client device 125. In some embodiments, content publisher computing device 120 may include a server for providing (or streaming) video or game content.

[0028] Client device 125 may include a computing device configured to acquire and display primary content provided by content publisher computing device 120 and content provided by content provider computing device 115 (e.g., third-party content items such as text, software programs, images, and / or videos). The client device may request and receive such content via network 105. Client device 125 may include a desktop computer, laptop computer, tablet device, smartphone, personal digital assistant, mobile device, consumer computing device, server, digital video recorder, set-top box, smart TV, video game console, or any other computing device capable of communicating via network 105 and consuming media content. Although Figure 1 A single client device 125 is shown, but environment 100 may include multiple client devices 125 served by data processing system 110.

[0029] While users of client device 125 can choose which primary content to access, the client device may have little control over the content provided by content provider computing device 115, as such content is typically selected automatically by data processing system 110. Third-party content providers or their corresponding content provider computing devices 115 may expose client device 125 to inappropriate or unwanted content, data privacy violations, cybersecurity threats, and other risks. To protect client device 125 from such risks and maintain the integrity of third-party content distribution, data processing system 110 can set policies for third-party content providers and / or their corresponding content provider computing devices 115 to adhere to. Data processing system 110 may also employ mechanisms to enforce these policies, for example, by detecting policy violations or policy-related flags and preventing the distribution of content associated with policy violators or policy flags. Once a policy violation or policy flag associated with an asset is detected, data processing system 110 can identify all assets associated with the source of the policy violation or the source of the asset associated with the policy flag and label these assets as, for example, malicious, suspicious, or blocked.

[0030] Data processing system 110 may include at least one computer server having one or more processors and a memory storing computer-executable instructions. For example, data processing system 110 may include multiple computer servers located in at least one data center or server cluster. In some embodiments, data processing system 110 may include a third-party content delivery system, such as an advertising server or advertising delivery system. Data processing system 110 may include at least one violation detection module 130, at least one asset clustering module 135, at least one policy classifier module 140, and at least one database 145. Each of the violation detection module 130, asset clustering module 135, and policy classifier module 140 may be implemented as a software module, a hardware module, or a combination of both. For example, each of these modules may include a processing unit, server, virtual server, circuitry, engine, agent, electrical appliance, or other logical device, such as a programmable logic array configured to communicate with database 145 or other computing devices via network 105. The computer-executable instructions of the data processing system 110 may include instructions that, when executed by one or more processors, cause the data processing system 110 to perform the operations discussed below regarding the violation detection module 130, the asset clustering module 135, the policy classifier module 140, or combinations thereof.

[0031] Database 145 can maintain a node network comprising multiple nodes and edges connecting corresponding node pairs. Each of the multiple nodes can represent a corresponding asset from multiple assets of multiple content providers or content sources. The database can use one or more data structures (such as trees, linked lists, tables, strings, or combinations thereof) to maintain the node network. The database can use the node network (or one or more data structures) to track assets used when distributing content to client device 125. The node network can be a heterogeneous network comprising nodes of different types or nodes corresponding to different types of assets. For example, the multiple assets represented by the nodes of the node network can include at least one asset of a first asset type and at least one asset of a second asset type different from the first asset type. The database can maintain data structures indicating policies and corresponding tags.

[0032] refer to Figure 2 The diagram illustrates an example network of nodes 200. Node network 200 is a heterogeneous network comprising nodes corresponding to different types of assets. For example, circle nodes 202a-202e represent content provider accounts. Diamond nodes 204a-204h represent different domains. Hexagonal nodes represent login pages, and square nodes represent data files of a given type (e.g., video files, text files, or image files). Pentagonal nodes represent payment information. Each link connecting a pair of nodes in node network 200 can represent some kind of relationship between the pair. For example, a link between a hexagonal node and a square node can indicate that the data file corresponding to the square node is provided by or can be downloaded from the login page represented by the hexagonal node. A link between a circle node—such as one of nodes 202a-202e—and a diamond node—such as any of nodes 204a-204h—can indicate that the domain represented by the diamond node belongs to the content provider account represented by the circle node, or provides content associated with that content provider account. The link between the circular and square nodes indicates that the data file represented by the square node belongs to the content provider account corresponding to the circular node. Finally, the link between the pentagonal nodes 206a-206c and the circular nodes 202a-202e indicates the payment information for each content provider account.

[0033] Typically, the edges of a node network, such as node network 200, may include edges between pairs of nodes corresponding to a website domain and a content provider account, respectively. The edges may indicate that at least one content item associated with the content provider account includes a link referencing an information resource (e.g., a webpage) associated with the website domain. Node network 200 may include edges between pairs of nodes corresponding to a second content provider account and payment information associated with that content provider account, respectively. The payment information may include information related to a bank account, billing account, or online invoicing account, or an identifier used to collect costs for service content associated with the content provider account. Node network 200 may include edges between pairs of nodes representing a website domain and an IP address associated with the website domain, respectively. The IP address may be the address of a host or server associated with the website domain. Node network 200 may include edges between pairs of nodes, including a node representing the website domain and another node representing an information resource associated with the website domain. A webpage (e.g., a login page) may be a page provided by the website domain.

[0034] Figure 2 The network of nodes 200 is provided for illustrative purposes and should not be construed as limiting. For example, the network of nodes 200 may have a larger number and other types of nodes. For instance, the network of nodes 200 may include other types of nodes, such as nodes corresponding to IP addresses, various types of resources, or hosts. Furthermore, the node network 200 may have more than... Figure 2 The number of links shown may be more or less. The network of nodes 200 can be dynamic, as any content provider can add new assets or remove existing assets over time. Furthermore, new relationships or correlations between existing nodes can be added, for example, by the content provider computing device 115, or discovered by the data processing system 110. Therefore, the data processing system 110 can update the network of nodes 200, for example, by adding new nodes, adding new links, removing existing nodes, or removing existing links.

[0035] The violation detection module 130 can detect violations or policy flags of the data processing system 110's policies, for example, through third-party content providers, their corresponding content provider computing devices 115, or their respective assets. The policy violation module 130 may rely on feedback from client device 125 or other computing devices (e.g., computing devices associated with the data processing system 110) reporting malicious, deceptive, or other conduct that does not comply with one or more policies associated with a given asset. Malicious, deceptive, or other unacceptable conduct may include, for example, the distribution of malware, stealthy redirects or sleight of hand, and other unacceptable practices by third-party content providers or their respective assets. Content items, login pages, or other resources from malicious content sources (e.g., third-party content providers or their respective hosts or domains) may cause malware to be downloaded to client device 125 accessing (or attempting to access) content items, login pages, or other resources.

[0036] A deceptive redirect occurs when a login page is configured to automatically redirect client device 125 to an alternative page that client device 125 does not intend to access. For example, client device 125 may receive content items from data processing system 110 that include links to login pages related to the topic of the content items. However, upon interaction with the links, the login page may redirect client device 125 to an alternative page, for example, associated with offensive or inappropriate content. Spoofing is the practice of presenting different content or URLs to client device 125 and computing devices associated with a content distribution system (e.g., data processing system 110 or a search engine). The host, website, or webpage of a malicious content source may include executable instructions to check whether the IP address of a given computing device is associated with client device 125 or the content distribution system, and determine which content or URL to serve to the computing device based on the result of the check. In this way, the host, website, or webpage may present client device 125 with content different from what is declared to data processing system 110.

[0037] Testers associated with data processing system 110 can test various assets against any practices or activities associated with policy tags of data processing system 110's policies and report any related policy tags or violations to violation detection module 130. In some implementations, violation detection module 130 can automatically detect related policy tags (e.g., policy tags applied to a given asset). For example, violation detection module 130 can check whether a website or login page associated with a content provider contains malware. Asset clustering module 135 can identify attribute sets that are instructions to perform deceptive redirects or check IP addresses to perform spoofing. Assets containing such software instructions are capable of violating data processing system 110's policies, even if they have not yet performed deceptive redirects or spoofing. Upon detecting a policy tag associated with a given asset, violation detection module 130 can provide asset clustering module 135 with an indication of that asset.

[0038] The asset clustering module 135 can identify nodes in the node network 200 that are associated with assets that have (or are mapped to) a policy tag (e.g., involve or are capable of violating a policy), referred to herein as seed nodes. The identified node (or seed node) can be a node representing an asset identified as associated with a policy tag, or, for example, another node corresponding to an asset of a given type associated with an asset identified as associated with a policy tag. For example, returning to the reference... Figure 2 The node indicated by the white arrow can represent an asset identified as associated with a policy tag. Here, an asset identified as associated with a policy tag can be, for example, a data file including malware files. However, the asset clustering module 135 can identify node 204b, indicated by the gray arrow, which represents a website domain that has some relationship (or correlation) with the asset identified as associated with the policy tag. For example, in this case, the data file identified as associated with the policy tag could be provided through a login page (or webpage) that is part of the website domain represented by node 204b.

[0039] The asset clustering module 135 may be interested in assets of a specific type (e.g., website domains or hosts, and other asset types) that are associated with assets identified as being associated with policy tags. Specifically, the asset clustering module 135 may be configured to cluster assets of predefined types and may begin the clustering process at a seed node corresponding to an asset of a predefined type that has some relationship (or correlation) with assets identified as being associated with policy tags. Performing asset clustering among a subset of nodes in the node network 200 (e.g., associated with predefined asset types) can significantly reduce the computational cost and time of the asset clustering process. Traversing all nodes in a large node network (e.g., with millions of nodes) can be computationally inefficient and may cause delays in the asset clustering process. In some implementations, the asset clustering module 135 may begin the asset clustering process at a node corresponding to an asset identified as being directly associated with a policy tag and may traverse all nodes in the node network 200.

[0040] Each asset can have its own identifier, and each node in the node network 200 can include (e.g., as metadata or as an identifier of the node itself) the identifier of the corresponding asset. The asset clustering module 135 can search the node network 200 for nodes with identifiers of assets identified as being associated with a policy tag. In some implementations, the database 145 can include a data structure that maps asset identifiers to identifiers of corresponding nodes in the node network 200. The asset clustering module 135 can use such a data structure to locate nodes within the node network 200 corresponding to assets identified as being directly associated with a policy tag. Once the asset clustering module 135 identifies a node corresponding to an asset identified as being directly associated with a policy tag, it can use links connected to that node to identify another node corresponding to an asset of a predefined type that is related (or associated with) the asset identified as being directly associated with a policy tag.

[0041] The asset clustering module 135 can identify a set (or combination) of two or more attributes of a seed node (or corresponding asset) identified as directly associated with a policy tag. The asset clustering module 135 can identify sets of attributes used to identify other assets belonging to the same owner or content source as the asset (or corresponding seed node) identified as directly associated with a policy tag. Generally, the more attributes used to cluster an asset or corresponding node, the higher the expected clustering accuracy. For example, using a single attribute is often insufficient to produce high-precision clusters. As an example, the fact that two domains are connected to the same content provider account, the same IP address, or the same resource does not necessarily mean that they belong to the same content source (or the participant behind the activity or behavior associated with the policy tag). For example, two otherwise unrelated website domains can share the same IP address through virtual hosting (e.g., they can be hosted by the same physical server in the cloud). Furthermore, different content provider accounts can cause corresponding content items to land on the same open-source content sharing platform (e.g., YouTube). Additionally, a given resource can be loaded by other unrelated domains, for example, because the websites hosted by the domain happen to be created using the same content management system (e.g., WordPress). However, when a group of domains (or other assets) or corresponding nodes share a combination of two or more independent attributes, this is often a strong indication that the domains (or other assets) were created, provided, or used by the same content source (e.g., a content provider entity or its agents).

[0042] The asset clustering module 135 can identify attribute sets using information or data from the node network 200. Specifically, the asset clustering module 135 can identify attribute sets based on links to or adjacent nodes of a seed node within the node network 200. The asset clustering module 135 can identify attributes using adjacent nodes within the node network 200 that are a few hops (e.g., two hops) away from the seed node. The asset clustering module 135 can identify attributes using metadata (if any) associated with the seed node in the node network 200. For example, attributes of a website domain may include content provider accounts associated with that website domain, payment information, login pages, data files, or any combination thereof. Attributes of a content provider account may include website domains (or more) associated with that content provider account, payment information, login pages, data files, or any combination thereof. For example, the asset clustering module 135 can identify attributes based on a predefined set of attribute types that have been previously tested and shown to provide accurate clustering results.

[0043] The asset clustering module 135 can calculate corresponding Point-to-Point Information (PMI) scores for subsets of nodes in the node network 200. The PMI score can indicate the likelihood that nodes within a subset sharing a set (or combination) of two or more attributes are associated with a common content source (e.g., a common third-party content provider). A subset of nodes can be defined, for example, as nodes in the node network 200 that have the same type as the seed node (e.g., nodes corresponding to a website domain or a host, and other types of nodes). In some implementations, a subset of nodes can be defined in other ways (e.g., not based on a predefined asset type). Nodes within a subset of nodes may share a set of attributes unexpectedly (e.g., randomly), or by design if such nodes belong to a common content source (e.g., a third-party content provider and / or a corresponding agent). The asset clustering module 135 can use the PMI score to distinguish between a set of attributes randomly shared by a group of nodes (or corresponding assets) due to association with a single content source and a set of nodes sharing the same set of attributes. The asset clustering module 135 can use information provided by the node network 200—such as links between node pairs—to calculate the PMI score for a subset of nodes. For example, the asset clustering module 135 can consider, for a given node, immediately adjacent nodes, adjacent nodes within the node network 200 that are a few (e.g., two) hops apart, or both, to determine whether an attribute is associated with the given node. In some implementations, the asset clustering module 135 can consider metadata associated with a given node to determine whether an attribute is associated with the given node.

[0044] Considering a set of two attributes, including attribute A and attribute B, the asset clustering module 135 can calculate the PMI score of a subset of nodes as follows:

[0045]

[0046] Where P(A) represents the probability that attribute A is associated with a node within a node subset. P(B) represents the probability that attribute B is associated with a node within a node subset. P(A,B) represents the joint probability that attributes A and B are associated with nodes within a node subset. The asset clustering module 135 can empirically calculate probabilities P(A), P(B), and P(A,B). For example, let attribute A indicate a website domain associated with a given payment information. If both the website domain and the payment information are associated with the same content provider account, then the website domain can be considered associated with the payment information. Attribute B may indicate a website domain that serves or provides a given data file. The asset clustering module 135 can empirically estimate P(A) as the number of domains associated with the payment information (or the number of corresponding nodes in the node subset) divided by the total number of domains (or the total number of nodes in the node subset). Similarly, the asset clustering module 135 can empirically estimate P(B) as the number of domains associated with the data file (or the number of corresponding nodes in the node subset) divided by the total number of domains (or the total number of nodes in the node subset). In addition, the asset clustering module 135 can empirically estimate P(A,B) as the number of domains associated with both payment information and data files (or the number of corresponding nodes in the node subset) divided by the total number of domains (or the total number of nodes in the node subset).

[0047] Equation (1) can also be written as:

[0048]

[0049] Where P(A / B) represents the conditional probability that attribute A is associated with a node given that attribute B is already associated with a node in the node subset, and P(B / A) represents the conditional probability that attribute B is associated with a node given that attribute A is already associated with a node.

[0050] For a set of n attributes X1, X2, ..., X... n Given a set (or combination) of nodes, the asset clustering module 135 can calculate the PMI score of the node subset as follows:

[0051]

[0052] Where, P(X1,…,X) n ) represents all the attributes X1,…,X n The first joint probability associated with nodes in the node subset, and P(X1,…,X) j-1 ,X j+1 ,…,X n ) indicates that, except for attribute X j All attributes other than X1,…,X n The second joint probability associated with the same node. Note that n represents a positive integer.

[0053] In some implementations, the asset clustering module 135 can assign different weights to different attributes when calculating the PMI score. For example, the asset clustering module 135 can use a weighted version of the PMI score in equation (1), as follows:

[0054]

[0055] Here, αA and αB represent the weighted values ​​of attributes A and B, respectively. Using different weighted values ​​for different attributes allows for assigning different rankings or different levels of importance to different attributes when calculating PMI scores.

[0056] In some implementations, the PMI score can be calculated using a function on the right-hand side of (1), (2), (3), or (4), or their respective approximations. For example, PMI(A,B) can be calculated using the following equation:

[0057]

[0058] Where f(·) denotes a real function. In one example, f(·) could be a logarithmic function. Therefore,

[0059]

[0060] Where log represents the logarithmic function with a known base (e.g., 2). Alternatively, PMI scores can be weighted according to the probability of the attribute. For example, PMI(A,B) can be calculated using the following equation:

[0061]

[0062] In some examples, where f(·) is a logarithmic function, we have

[0063]

[0064] Similar extensions can be made to the PMI scores defined in (3) with n attributes or their weighted versions in (4). Further note that, for computational simplicity, approximate or fixed-point calculations of the PMI scores can be used in practice.

[0065] The asset clustering module 135 can use PMI scores associated with a subset of nodes to identify node clusters in a heterogeneous network that include a set (or combination) of two or more attributes. Specifically, the asset clustering module 135 can determine whether nodes sharing a set of attributes (within the node subset) represent an asset cluster associated with a single content source (e.g., a single content provider and / or corresponding agent) based on the value of the PMI score of the node subset relative to the set of attributes. The asset clustering module 135 can compare the PMI score to a predefined threshold. If the PMI score exceeds (or is greater than or equal to) the predefined threshold, the asset clustering module 135 can identify all nodes within the node subset that share a set of attributes (or are associated with the set of attributes) and form a cluster consisting of these nodes. In some implementations, the threshold can be equal to 1. In some implementations, the asset clustering module 135 or the data processing system 110 can test (e.g., based on training or test data) multiple thresholds and select, for example, a threshold that provides the best clustering precision-recall tradeoff. The threshold used may depend on, for example, the PMI formula used by the asset clustering module 135 (such as the PMI formula described in equations (1)-(4)).

[0066] A relatively high PMI value indicates a hidden relationship (or hidden design) between nodes or corresponding assets that share a set of attributes within a subset of nodes. Specifically, a relatively high PMI value indicates that nodes or corresponding assets that share a set of attributes within a subset of nodes are most likely to belong to the same content source.

[0067] refer to Figure 3 The diagram illustrates the clustering results within the illustrated node network 200. In this case, and as previously stated, the asset clustering module 135 defines a subset of nodes as hexagonal nodes 204a-204h representing domain assets (in...). Figure 3 (Shown in gray in the middle). Using PMI scores, for example, for attributes indicating association with content provider accounts 202a and 202b (both associated with seed node 204B) and association with a login page that provides a given data file (in this case, the data file is identified as being directly associated with a policy tag or policy violation), the asset clustering module 135 can identify cluster 302 as belonging to (or associated with) a single content source.

[0068] In some implementations, the cluster may include nodes from a subset of nodes as well as other nodes in the node network 200. For example, the asset clustering module 135 can identify groups of nodes with a set of attributes within a subset of nodes (e.g., nodes corresponding to a domain). Figure 3In the example, the asset clustering module 135 can first identify a set of nodes 204a-204c. Then, the asset clustering module 135 can identify, for example, content provider accounts associated with the identified set of nodes that have (or share) that set of attributes. The asset clustering module 135 can then identify all nodes (or assets) associated with the identified content provider nodes as nodes forming a cluster. For example, in... Figure 3 In the example shown, the asset clustering module 135 can identify content provider accounts 202a and 202b as associated with domains 204a-204d, and then identify cluster 302 as all nodes or corresponding assets in content provider accounts 202a and 202b. In some implementations, the asset clustering module 135 may employ different methods to identify all nodes (or assets) of different types in a cluster. Clusters of assets include those identified as being directly associated with policy tags.

[0069] Based on a first asset identified as directly associated with a policy tag of the content distribution system's policy, the policy classifier module 140 can store the associations between node clusters and tags identified by the asset clustering module 135 in one or more data structures. The policy classifier module 140 can use tags to classify a first set of assets among multiple assets corresponding to node cluster 200. For example, the policy classifier module 140 can use tags to classify assets or sets of assets into rogue, malicious, suspicious, or blocked categories, among others. In other words, since a cluster of nodes (e.g., cluster 302) represents assets corresponding to a single content source and including assets identified as associated with a policy tag, a cluster can be considered as representing assets belonging to an entity (or associated with) a source of activity or behavior that violates at least one policy of the data processing system 110 or is the subject of a policy tag. Therefore, the policy classifier module 140 can label the entire cluster (or its assets) as suspicious, rogue, untrustworthy, or malicious, among other categories. The policy classifier module 140 can prevent or restrict the provision (or distribution) of one or more assets of the cluster (or tagged assets), such as data files, login pages, or content items, to the client device 135. The policy classifier module 140 can prevent or restrict one or more assets of the cluster from participating in any content distribution activity. For example, the policy classifier module 140 can prevent or restrict content provider accounts in the cluster from participating in any auctions that provide third-party content to the client device 135.

[0070] PMI-based clustering methods can utilize both local and global information. Locally, the data processing system 110 can consider the set of common neighbors of a set of nodes in a node network (e.g., to define attributes). Globally, the data processing system 110 can use the total number of nodes (e.g., domains) to calculate the probabilities constituting a PMI score. Note that clusters can be ranked using local information alone, since the product PMI×n does not depend on, for example, the total number n of domains shared by all clusters. However, the threshold for high-precision clustering will still depend on n or some other global information (e.g., PMI score quantiles). Since the proposed clustering method is primarily a local method, it allows for efficient cluster updates when the node network is modified, e.g., slightly modified. In contrast, global methods must recompile all clusters from scratch even when a single node or relationship changes. The proposed clustering method is also easier to implement on distributed systems, as it typically does not require a large amount of shared memory. Because it relies primarily on local information, the above-described PMI-based clustering is naturally suitable for online implementation. Specifically, as new entities and relationships are added to the graph in real time, the size of the graph (or node network) used to formulate or calculate PMI scores and the attributes or links associated with its nodes can be easily updated.

[0071] refer to Figure 4 The diagram illustrates a flowchart of a method 400 for implementing a strategy associated with content distribution. Method 400 may include maintaining a network of nodes representing multiple assets (step 405). Method 400 may include detecting that a first asset among the multiple assets has a policy tag of the strategy or is associated with a policy tag of the strategy (step 410), and identifying a first node in the node network associated with the first asset (step 415). Method 400 may include identifying multiple attributes of the first node (step 420), and calculating a corresponding PMI score for a subset of nodes in the node network, the PMI score indicating the likelihood that a node with multiple attributes is associated with a single content source (step 425). Method 400 may include identifying clusters of nodes with multiple attributes (step 430), and storing the association between clusters of nodes and tags used to classify the asset set (step 435).

[0072] refer to Figure 1-4 Method 400 can be found in the above reference. Figure 1-3 The data processing system 110 discussed here operates. For example, the node network can be a heterogeneous network of nodes, such as node network 200, where the corresponding nodes represent at least two different types of assets. Assets can be associated with, for example, multiple third-party content providers. Policy tags detected in association with the first asset can indicate violations of the policies of data processing system 110, such as classifications of related content distribution, asset activities, or asset behaviors.

[0073] The above description of PMI-based clustering, for example, in applications such as identifying assets associated with a single content source through a content distribution system, illustrates an application of PMI-based clustering. However, the aforementioned clustering methods (or more) can generally be applied to classification problems where entities can be of multiple types, and the relationships between entities can naturally be represented as a graph (or a network of nodes). This is particularly useful when baseline ground truth class labels are unavailable or difficult to obtain, such as in unsupervised or semi-supervised settings. In such cases, the high-precision clusters generated by the aforementioned PMI-based clustering methods (or more) can be used as a replacement for the baseline ground truth data. In semi-supervised settings, PMI-based clustering allows for the extrapolation of several existing baseline ground truth labels to the entire cluster. Note that due to the high precision of the clusters, PMI-based clustering does not require labeling most examples in the cluster. Instead, a single labeled example may be sufficient to extrapolate to the entire cluster. Therefore, PMI-based clustering is naturally well-suited for classification settings with a small sample size. In unsupervised settings, high-precision clustering can be used to efficiently guide the manual labeling process. In many applications, including those described above regarding identifying potential rogue or suspicious assets in content distribution systems, reviewing and tagging the entire cluster of relevant entities is easier than reviewing and tagging each entity individually. The context provided by high-precision clustering makes the manual review and tagging process more efficient and robust.

[0074] Applications particularly well-suited for the PMI-based clustering described above include abuse detection, such as payment fraud, spam, and organized crime groups. Abuse networks tend to possess a degree of redundancy, making them more resilient to implementation. The PMI-based clustering described above can derive high-precision features from this redundancy. In abuse networks, nodes can represent publishers, HTTP cookies, advertising campaigns, or combinations thereof. Links in abuse networks can represent publisher-cookie relationships, publisher-campaign relationships, and / or cookie-campaign relationships.

[0075] Another example application is anomaly detection on websites, which requires grouping websites and identifying anomalous patterns. Nodes in the node network used for this application can represent domain names, IP addresses, registration information from WhoIs Domain Lookup, or combinations thereof. Node attributes can include web page resources (e.g., images, text, etc.), registration time, domain access counts (if available), or combinations thereof. Links in the node network can represent relationships between domains and IP addresses, between domains and registrant information, between domains and other domains, or combinations thereof.

[0076] Another exemplary application is financial risk prediction. In such an application, given a high-precision cluster of similar stocks and only one or a few examples of defaulting stocks within each cluster, the risk of stock loss can be predicted with high accuracy. The node network in this application can include nodes representing stock symbols and news websites, etc. Node attributes can include industry, quarterly reports, news report summaries, news report age, or combinations thereof, etc. Links in the node network can represent (stock symbol, stock symbol) pairs and (stock symbol, news website) pairs.

[0077] Another example application of PMI-based clustering is social network analysis, where the goal is to identify self-organizing communities within large-scale social networks. In this context, nodes can represent user accounts, client devices accessing the user account from them, IP addresses used to access the user account, posts shared and viewed by the user account, posts viewed by the user account, or combinations thereof. Links can represent friendship relationships between user accounts, relationships between user accounts and client devices and / or IP addresses accessing the user account from them, relationships between user accounts and posts viewed or shared by that account, or combinations thereof.

[0078] Another application of PMI-based clustering is insurance forecasting, where the goal is to estimate the likelihood and number of claims given a high-precision cluster of insured property and a small number of examples from each cluster from which claims have been made.

[0079] Another application of PMI-based clustering is epidemic prediction. High-precision clustering of susceptible individuals can allow for accurate targeting of vaccination programs in the initial stages, when only a small number of new epidemic cases are known. The online nature of the proposed clustering method is particularly useful in this application.

[0080] Another application of PMI-based clustering is similar audience targeting. The concept of similar audiences plays a central role in advertising, online marketing, and entertainment. High-precision clustering allows for the expansion of a small number of successfully targeted examples to a larger audience.

[0081] Another application of PMI-based clustering is the discovery of protein complexes in protein-protein interaction networks. High-precision clustering allows the generalization of known interactions of only a few proteins to the entire cluster.

[0082] For systems discussed here that collect or utilize personal information about users, users can be given the opportunity to control whether a program or feature can collect personal information (e.g., information about the user's social networks, social actions or activities, user preferences, or the user's current location), or to control whether or how content that may be more relevant to the user is received from the content server. Additionally, some data can be processed in one or more ways before storage or use, such that certain information about the user is removed when generating parameters (e.g., demographic parameters). For example, a user's identity can be processed so that the user's identifying information cannot be determined, or the user's geographic location can be generalized (e.g., to the city, zip code, or state level) when location information is available, making it impossible to determine the user's specific location. Therefore, users can control how content servers collect and use information about them.

[0083] Figure 5 The illustration shows a general architecture of a computer system 500 according to some implementations. This computer system 500 can be used to implement any computer system discussed herein, including system 110 and its components such as violation detection module 130, asset clustering module 135, and policy classifier module 140. The computer system 500 can be used to provide information for display via network 105. Figure 5 The computer system 500 includes one or more processors 520 communicatively coupled to memory 525, one or more communication interfaces 505, and one or more output devices 510 (e.g., one or more display units) and one or more input devices 515. The processors 520 may be included in the data processing system 110 or other components of the system 110, such as a violation detection module 130, an asset clustering module 135, and a policy classifier module 140.

[0084] exist Figure 5 In the computer system 500, the memory 625 may include any computer-readable storage medium and may store computer instructions, such as processor-executable instructions, for implementing the various functions described herein for the various systems, and any data associated with, generated therefrom, or received via a communication interface or input device (if present). See again Figure 1 In the environment 100, the data processing system 110 may include a memory 525 to store data structures and / or information related to, for example, node networks or PMI scores. The memory 525 may include a database 145. Figure 5 The processor(s) 520 shown can be used to execute instructions stored in memory 625, and in doing so, can also read from or write to memory various information processed and / or generated according to the execution of instructions.

[0085] Figure 5 The processor 520 of the computer system 500 shown can also be communicatively coupled to or control the communication interface 505 to send or receive various information according to the execution of instructions. For example, the communication interface 505 can be coupled to a wired or wireless network, a bus, or other communication device, and thus allow the computer system 500 to send information to or receive information from other devices (e.g., other computer systems). Although not explicitly stated in the text... Figure 1 The system is explicitly shown, but one or more communication interfaces facilitate the flow of information between components of system 500. In some embodiments, the communication interface may be configured (e.g., via various hardware or software components) to provide a website as an access portal to at least some aspects of computer system 500. Examples of communication interface 505 include a user interface (e.g., a webpage) through which a user can communicate with data processing system 110.

[0086] For example, it can provide Figure 5 The output device 510 of the computer system 500 shown allows for viewing or otherwise perceiving various information in conjunction with the execution of instructions. For example, an input device 515 may be provided to allow a user to manually adjust, select, input data, or interact with the processor in any of these ways during instruction execution. This document further provides additional information relating to general computer system architectures that can be used in the various systems discussed herein.

[0087] The implementation of the subject matter and operations described in this specification can be implemented in digital electronic circuits or in computer software contained on tangible media, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by or control of the operation of a data processing device. The program instructions can be encoded on artificially generated propagated signals, such as machine-generated electrical, optical, or electromagnetic signals, generated to encode information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium can be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof, or is included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof. Furthermore, although the computer storage medium is not a propagated signal, it can include a source or destination of computer program instructions encoded in artificially generated propagated signals. Computer storage media may also be one or more separate physical components or media (e.g., multiple CDs, disks or other storage devices), or may be included in one or more separate physical components or media (e.g., multiple CDs, disks or other storage devices).

[0088] The features disclosed herein can be implemented on a smart TV module (or connected TV module, hybrid TV module, etc.) that may include a processing module configured to integrate internet connectivity with more traditional television program sources (e.g., received via cable, satellite, over-the-air, or other signals). The smart TV module may be physically incorporated into a television set or may include a separate device such as a set-top box, Blu-ray or other digital media player, game console, hotel TV system, and other related devices. The smart TV module may be configured to allow viewers to search for and find videos, movies, photos, and other content on the internet, local cable TV channels, satellite TV channels, or stored on a local hard drive. A set-top box (STB) or set-top box unit (STU) may include an information appliance that may contain a tuner and connect to the television set and external signal sources, converting the signals into content and then displaying that content on the television screen or other display devices. The smart TV module may be configured to provide a main screen or top-level screen that includes icons for multiple different applications, such as web browsers and multiple streaming services, connected cable or satellite media sources, other web “channels,” etc. The smart TV module can also be configured to provide users with an electronic program guide. Companion applications for the smart TV module can operate on mobile computing devices to provide users with additional information about available programming, allow users to control the smart TV module, etc. In alternative implementations, these features can be implemented on laptops or other personal computers, smartphones, other mobile phones, handheld computers, tablet PCs, or other computing devices.

[0089] The operations described in this specification can be implemented as operations performed by a data processing device on data stored on one or more computer-readable storage devices or received from other sources.

[0090] The terms "data processing apparatus," "data processing system," "user equipment," or "computing device" include all types of apparatus, devices, and machines for processing data, including, for example, programmable processors, computers, systems-on-a-chip, or a combination thereof. The apparatus may include special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. The apparatus and execution environment can implement various computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures. Content request module 130 and content selection module 135 may include or share one or more data processing apparatuses, computing devices, or processors.

[0091] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, declarative or procedural languages. Computer programs can be deployed in any form, including as standalone programs or as modules, components, subroutines, objects, or other units suitable for use in a computing environment. A computer program may, but does not necessarily, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or code sections). A computer program can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communication network.

[0092] The processes and logic flows described in this specification can be executed by one or more programmable processors that execute one or more computer programs to perform actions by manipulating input data and generating outputs. The processes and logic flows can also be executed by special-purpose logic circuitry (e.g., FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits)), and the apparatus can also be implemented as special-purpose logic circuitry (e.g., FPGAs or ASICs).

[0093] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for performing actions according to instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from or transfer data to or from them, or both. However, a computer does not need to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive). Devices suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM discs. Processors and memory can be supplemented by dedicated logic circuits or incorporated into dedicated logic circuits.

[0094] To provide interaction with the user, the implementation of the subject matter described in this specification can be carried out on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube), plasma, or LCD (liquid crystal display) monitor, and a keyboard and pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can include any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a webpage to a web browser on the user's client device in response to a request received from a web browser.

[0095] The implementations of the subjects described in this specification can be implemented in a computing system that includes backend components, such as a data server, or middleware components, such as an application server, or frontend components, such as a client computer with a graphical user interface or a web browser through which a user can interact with the implementations of the subjects described in this specification, or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (“LANs”) and wide area networks (“WANs”), interconnected networks (e.g., the Internet) and peer-to-peer networks (e.g., self-organizing peer-to-peer networks).

[0096] Computing system 500 may include a content provider computing device 115, a publisher computing device 120, a client device 125, or a server or computing device of data processing system 110. For example, data processing system 110 may include one or more servers in one or more data centers or server clusters. Clients and servers are typically geographically distant from each other and typically interact via a communication network. The client-server relationship is established by means of computer programs running on respective computers and having a client-server relationship with each other. In some implementations, the server sends data (e.g., HTML pages) to the client device (e.g., for the purpose of displaying data to a user interacting with the client device and receiving user input from the user interacting with the client device). Data generated at the client device (e.g., the result of user interaction) may be received at the server from the client device.

[0097] While this specification contains numerous specific implementation details, these details should not be construed as limiting any invention or potentially claimed scope, but rather as descriptions of features specific to particular embodiments of the systems and methods described herein. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be removed from the combination, and the claimed combination may involve sub-combinations or variations thereof.

[0098] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring that they be performed in the specific order shown or sequentially, or that all of the shown operations be performed to achieve the desired result. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired result.

[0099] In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. For example, content request module 130 and content selection module 135 may be part of data processing system 110, a single module, a logical device with one or more processing modules, or part of one or more servers or search engines.

[0100] Some illustrative embodiments and implementations have now been described. It is obvious that the foregoing is illustrative and not limiting, and has been presented by way of example. In particular, although many of the examples presented herein involve specific combinations of method actions or system elements, those actions and elements can be combined in other ways to achieve the same purpose. Actions, elements, and features discussed in connection with only one embodiment are not intended to exclude similar roles in other embodiments or implementations.

[0101] The wording and terminology used herein are for descriptive purposes and should not be considered limiting. The use of “comprising,” “including,” “having,” “containing,” “involving,” “characterized in,” “featured in,” and variations thereof throughout this document is intended to cover the items listed thereafter, their equivalents and additional items, as well as alternative implementations consisting of the items listed thereafter. In one implementation, the systems and methods described herein consist of one, more than one, each combination of, or all of the described elements, actions, or components.

[0102] Any reference to an implementation of a system or method, element, or action mentioned herein in the singular may include implementations that include multiple such elements, and any plural reference to any implementation, element, or action herein may include implementations that include only a single element. References in the singular or plural form are not intended to limit the currently disclosed systems or methods, their components, actions, or elements to a singular or plural configuration. A reference to any action or element based on any information, action, or element may include an action or element that is at least partially based on an implementation of that information, action, or element.

[0103] Any implementation disclosed herein may be combined with any other implementation, and references to “implementation,” “some implementations,” “alternative implementations,” “various implementations,” “one implementation,” etc., are not necessarily mutually exclusive and are intended to indicate that a particular feature, structure, or characteristic described in connection with that implementation may be included in at least one implementation. Such terms as used herein do not necessarily refer to the same implementation. Any implementation may be combined inclusively or exclusively with any other implementation in any manner consistent with the aspects and implementations disclosed herein.

[0104] A reference to "or" can be interpreted as inclusive, such that any term described using "or" can refer to any one, more than one, or all of the terms described.

[0105] Where reference numerals follow technical features in the drawings, detailed descriptions, or any claims, the sole purpose of including these reference numerals is to enhance the comprehensibility of the drawings, detailed descriptions, and claims. Therefore, reference numerals, and their absence, do not limit the scope of any claim element.

[0106] The systems and methods described herein may be implemented in other specific forms without departing from the characteristics of the invention. Although the examples provided herein relate to the display of content of information resources, the systems and methods described herein may be applicable to other environments. The foregoing embodiments are illustrative and not limiting of the systems and methods described. Therefore, the scope of the systems and methods described herein is indicated by the appended claims rather than the foregoing description, and modifications falling within the meaning and scope of equivalents of the claims are also included.

Claims

1. A system comprising: At least one processor; as well as Memory storing computer-executable instructions that, when executed by at least one processor, cause the at least one processor to: Maintain a heterogeneous network of nodes, which includes multiple nodes and edges connecting corresponding pairs of nodes. Each of the multiple nodes represents a corresponding asset among multiple assets corresponding to multiple content sources. The multiple assets include at least one asset of a first asset type and at least one asset of a second asset type. The first asset among multiple assets is identified as having a tag associated with the content distribution system's strategy; Identify the first node associated with the first asset in a heterogeneous network of nodes; Identify combinations of two or more attributes of the first node; For a subset of nodes in a heterogeneous network, calculate the corresponding inter-node mutual information (PMI) score. The PMI score is based on node pairs to indicate the likelihood that a node in the node subset with a combination of the two or more attributes is associated with a single content source. Using PMI scores associated with a subset of nodes to identify node clusters in a heterogeneous network that include combinations of the two or more attributes, wherein identifying node clusters includes: The corresponding PMI score of a subset of nodes is compared with a predefined threshold, wherein the predefined threshold is selected from multiple thresholds based on a PMI formula, and wherein the predefined threshold is selected from the multiple thresholds based on a clustering precision-recall tradeoff; and If the corresponding PMI score exceeds a predefined threshold, then nodes that include a combination of two or more attributes are determined to belong to the node cluster; and The association between the node cluster and tags is stored in one or more data structures, the tags being based on a first asset having a tag associated with a content distribution system strategy and used to classify a first asset set among multiple assets corresponding to the node cluster.

2. The system as claimed in claim 1, wherein, The computer-executable instructions, when executed by at least one processor, further cause the at least one processor to limit the provision of one or more assets from the first asset set.

3. The system as described in claim 1, wherein, The first asset type or the second asset type includes at least one of the following: content provider account, website domain, information resource, Internet Protocol IP address, data file, content item, or payment information.

4. The system as claimed in claim 1, wherein, When executed by at least one processor, the computer-executable instructions further enable the at least one processor to select each node in a subset of nodes based on the asset type of the first node.

5. The system as described in claim 4, wherein, The asset type of the first node corresponds to the website domain.

6. The system as claimed in claim 1, wherein, The PMI score, which indicates the probability that nodes in a subset of indicators with attributes A and B are associated with a single content source, is defined as... , where P(A) represents the probability that attribute A is associated with a node in the node subset, P(B) represents the probability that attribute B is associated with the node in the node subset, and P(A, B) represents the joint probability that attributes A and B are associated with the node in the node subset.

7. The system as claimed in claim 1, wherein, The indicator node subset has n attributes X1, X2, …, X n The PMI score for the probability that a combination of nodes is associated with a single content source is defined as... Where P(X1,…,X) n ) represents all attributes X1, …, X n The first joint probability associated with nodes in the said node subset, and P(X1,…, X…) j-1 ,X j+1 ,…,X n ) indicates that, in addition to attribute X j All attributes other than X1,…, X n The second joint probability associated with the node in the subset of nodes.

8. The system of claim 1, wherein, The predefined threshold is determined based on the quantiles of the PMI scores associated with the subset of nodes.

9. The system of claim 1, wherein, The edges connecting the corresponding node pairs include at least one of the following: The first side between the first pair of nodes, the first pair of nodes including a third node representing a first website domain and a fourth node representing a first content provider account, the first side indicating that at least one content item associated with the first content provider account includes a link referencing a first information resource associated with the first website domain; The second side between the second pair of nodes, the second pair of nodes includes a fifth node representing the second content provider account and a sixth node representing payment information associated with the second content provider account; The third edge between the third pair of nodes, the third pair of nodes including the seventh node representing the second website domain and the eighth node representing the Internet Protocol IP address associated with the second website domain; or The fourth edge between the fourth pair of nodes, the fourth pair of nodes including the ninth node representing the third website domain and the tenth node representing the second information resource associated with the third website domain.

10. The system of claim 1, wherein detecting that the first asset among the plurality of assets has the tag associated with the policy of the content distribution system includes detecting the distribution of malware or malicious content.

11. A method comprising: A heterogeneous network of nodes is maintained by a data processing system including one or more processors. The heterogeneous network of nodes includes multiple nodes and edges connecting corresponding pairs of nodes. Each of the multiple nodes represents a corresponding asset among multiple assets corresponding to multiple content sources. The multiple assets include at least one asset of a first asset type and at least one asset of a second asset type. The data processing system detects that the first asset among multiple assets has a tag associated with the content distribution system's strategy; The data processing system identifies the first node associated with the first asset in a heterogeneous network of nodes; The data processing system identifies combinations of two or more attributes of the first node; The data processing system calculates the corresponding inter-node mutual information (PMI) scores for a subset of nodes in a heterogeneous network. The PMI scores are based on node pairs and indicate the likelihood that a node in the node subset with a combination of the two or more attributes is associated with a single content source. The data processing system uses PMI scores associated with a subset of nodes to identify node clusters in a heterogeneous network that include combinations of two or more of the aforementioned attributes, wherein identifying node clusters includes: The corresponding PMI score of a subset of nodes is compared with a predefined threshold, wherein the predefined threshold is selected from multiple thresholds based on a PMI formula, and wherein the predefined threshold is selected from the multiple thresholds based on a clustering precision-recall tradeoff; and If the corresponding PMI score exceeds a predefined threshold, then nodes that include a combination of two or more attributes are determined to belong to the node cluster; and The data processing system stores the association between node clusters and tags in one or more data structures, the tags being based on a first asset having a tag associated with a content distribution system strategy and used to classify a first asset set among the plurality of assets corresponding to the node cluster.

12. The method of claim 11, further comprising: The data processing system limits the provision of one or more assets from the first asset set.

13. The method according to claim 11, wherein, The first or second asset type includes at least one of the following: content provider account, website domain, information resource, Internet Protocol IP address, data file, content item, or payment information.

14. The method of claim 11, further comprising: The data processing system selects each node in the node subset based on the asset type of the first node.

15. The method according to claim 11, wherein, The PMI score, which indicates the probability that nodes in a subset of indicators with attributes A and B are associated with a single content source, is defined as... , where P(A) represents the probability that attribute A is associated with a node in the node subset, P(B) represents the probability that attribute B is associated with the node in the node subset, and P(A, B) represents the joint probability that attributes A and B are associated with the node in the node subset.

16. The method according to claim 11, wherein, The indicator node subset has n attributes X1, X2, …, X n The PMI score for the probability that a combination of nodes is associated with a single content source is defined as... Where P(X1,…,X) n ) represents all attributes X1, …, X n The first joint probability associated with nodes in the said node subset, and P(X1,…, X…) j-1 ,X j+1 ,…,X n ) indicates that, in addition to attribute X j All attributes other than X1,…, X n The second joint probability associated with the node in the subset of nodes.

17. The method according to claim 11, wherein, The edges connecting the corresponding node pairs include at least one of the following: The first side between the first pair of nodes, the first pair of nodes including a third node representing a first website domain and a fourth node representing a first content provider account, the first side indicating that at least one content item associated with the first content provider account includes a link referencing a first information resource associated with the first website domain; The second side between the second pair of nodes, the second pair of nodes includes a fifth node representing the second content provider account and a sixth node representing payment information associated with the second content provider account; The third edge between the third pair of nodes, the third pair of nodes including the seventh node representing the second website domain and the eighth node representing the Internet Protocol IP address associated with the second website domain; or The fourth edge between the fourth pair of nodes, the fourth pair of nodes includes the ninth node representing the third website domain and the tenth node representing the second information resource associated with the third website domain.

18. The method according to claim 11, wherein, Detecting the first asset among multiple assets with a tag associated with the content distribution system's policy includes detecting the distribution of malware or malicious content.

19. A non-transitory computer-readable medium storing computer-executable instructions, which, when executed by at least one processor, cause the at least one processor to: Maintain a heterogeneous network of nodes, which includes multiple nodes and edges connecting corresponding pairs of nodes. Each of the multiple nodes represents a corresponding asset among multiple assets corresponding to multiple content sources. The multiple assets include at least one asset of a first asset type and at least one asset of a second asset type. The first asset among multiple assets is identified as having a tag associated with the content distribution system's strategy; Identify the first node associated with the first asset in a heterogeneous network of nodes; Identify combinations of two or more attributes of the first node; For a subset of nodes in a heterogeneous network, calculate the corresponding inter-node mutual information (PMI) score. The PMI score indicates the likelihood that a node in the subset of nodes with a combination of the two or more attributes is associated with a single content source, based on the node pairs. Using PMI scores associated with a subset of nodes to identify node clusters in a heterogeneous network that include combinations of the two or more attributes, wherein identifying node clusters includes: The corresponding PMI score of a subset of nodes is compared with a predefined threshold, wherein the predefined threshold is selected from multiple thresholds based on a PMI formula, and wherein the predefined threshold is selected from the multiple thresholds based on a clustering precision-recall tradeoff; and If the corresponding PMI score exceeds a predefined threshold, then nodes that include a combination of two or more attributes are determined to belong to the node cluster; and The association between the node cluster and tags is stored in one or more data structures, the tags being based on a first asset having a tag associated with a content distribution system strategy and used to classify a first asset set among multiple assets corresponding to the node cluster.

Citation Information

Patent Citations

  • Detecting malware infestations in large-scale networks

    US8959643B1