Systems and methods for providing recommendations based on seed supervised learning

CN110720099BActive Publication Date: 2026-09-29BEIJING DIDI INFINITY TECH & DEV CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201780091594.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2017-06-05
Publication Date
2026-09-29
Estimated Expiration
2037-06-05

AI Technical Summary

Technical Problem

[0004]然而,社交网络数据的大小通常很大,而现有分析社交网络数据的技术往往不够充分,效率低下,并且不能充分利用隐藏模式、用户行为、市场趋势、客户偏好以及嵌入其中的其他有用信息数据

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110720099B_ABST
    Figure CN110720099B_ABST
Patent Text Reader

Abstract

Systems and methods for providing recommendations based on seed supervised learning are provided. The method can include obtaining, over a communication network, similarity data associated with a first entity, a second entity, and a third entity, and obtaining, over the communication network, external data associated with the first entity and the second entity. The method can further include training a classification model based on the external data and the similarity data. The method can also include determining, based on the classification model, an expected score for the third entity, and providing, over the communication network, a recommendation based on the expected score to the third entity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to big data and machine learning technologies, and in particular to systems and methods for providing recommendations based on seed-supervised learning. Background Technology

[0002] Due to recent advancements in computer technology and the widespread adoption of the internet, many people have begun using platforms that leverage databases. Consequently, these databases often require "big data" analytics. Big data analytics is the process of examining "big data"—large or complex datasets—to reveal hidden patterns, user behavior, market trends, customer preferences, and other useful information. "Big data" is typically stored in database management systems (such as Oracle®, Teradata®, PostgreSQL, Microsoft SQL Server, and MySQL™ database management systems), which do not possess the capability to analyze these datasets.

[0003] Analyzing big data manually or even semi-automatically requires significant human resources. For example, a company might hire a team of engineers to develop solutions that provide intelligent suggestions to social network users in order to effectively acquire new customers at the lowest cost. These social networks may include both existing and potential customers. By using these social networks, one can understand the behavior of potential customers based on information relevant to existing customers.

[0004] However, social network data is typically very large, and existing technologies for analyzing social network data are often insufficient, inefficient, and fail to fully utilize hidden patterns, user behavior, market trends, customer preferences, and other useful information embedded within it.

[0005] Given these and other drawbacks and problems of big data analytics, machine learning is needed to improve the systems and methods for providing recommendations, and more specifically, seed-supervised learning is required. Summary of the Invention

[0006] One aspect of this application provides a system for providing recommendations to entities. The system may include a memory and a processor. The processor may be configured to acquire similarity data associated with a first entity, a second entity, and a third entity, and to acquire external data associated with the first entity and the second entity. The processor may be further configured to train a classification model based on the external data and the similarity data. The processor may also be configured to determine an expected score for the third entity based on the classification model, and to provide recommendations to the third entity based on the expected score.

[0007] Another aspect of this application provides a computer-implemented method for providing recommendations to entities. The method may include acquiring similarity data associated with a first entity, a second entity, and a third entity via a communication network, and acquiring external data associated with the first entity and the second entity via the communication network. A processor may be further configured to train a classification model based on the external data and the similarity data. The processor may also be configured to determine an expected score for the third entity based on the classification model, and provide recommendations based on the expected score to the third entity via the communication network.

[0008] Another aspect of this application provides a non-transitory computer-readable medium. The non-transitory computer-readable medium stores a set of instructions, which, when executed by at least one processor of a recommendation system, cause the recommendation system to perform a method for providing recommendations to entities. The method includes acquiring similarity data associated with a first entity, a second entity, and a third entity, and acquiring external data associated with the first entity and the second entity. The method may further include training a classification model based on the external data and the similarity data. The method may also include determining an expected score for the third entity based on the classification model; and providing recommendations to the third entity based on the expected score.

[0009] It should be understood that the foregoing general description and the following detailed description are merely exemplary and illustrative, and do not constitute a limitation on the specific application. Attached Figure Description

[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments and, together with the specification, serve to explain the principles of this application. In the drawings:

[0011] Figure 1 This is a schematic diagram of an exemplary system for providing recommendations based on seed-supervised learning, according to some embodiments of this application.

[0012] Figure 2 This is a block diagram of an exemplary system for providing recommendations based on seed-supervised learning, according to some embodiments of this application.

[0013] Figure 3 This is a diagram of an exemplary social network shown according to some embodiments of this application.

[0014] Figure 4 This is a flowchart of an exemplary process for providing recommendations based on seed-supervised learning.

[0015] Figure 5 This is a flowchart illustrating an exemplary process for training a classification model according to some embodiments of this application.

[0016] Figure 6 This is a flowchart illustrating an exemplary process for determining a seed for a conversion model, according to some embodiments of this application. Detailed Implementation

[0017] This application provides novel systems and methods for providing recommendations based on seed-supervised learning. Specifically, the disclosed systems and methods acquire new customers by providing intelligent recommendations to users of social networks through machine learning. Machine learning trains computers to learn, for example, to perform pattern recognition or artificial intelligence, without necessarily requiring specialized programming. Specifically, this application utilizes semi-supervised learning to predict features of a large amount of data based on a small number of known features from a small amount of data. The disclosed systems and methods improve existing computer systems by providing new systems and methods to train computer systems in novel ways. For example, one aspect of the disclosed embodiments provides new systems and methods to provide intelligent recommendations to a large number of users based on the characteristics of a small group of users. To provide these improvements, the disclosed systems and methods can be implemented using hardware, firmware and / or software, and combinations of dedicated hardware, firmware and / or software, such as specially constructed and / or programmed machines for performing the functions associated with the disclosed method steps. However, in some embodiments, the disclosed systems and methods can be implemented in dedicated electronic devices.

[0018] According to embodiments of this application, a system for providing recommendations to an entity may include a processor and a memory device storing instructions. As indicated in this application, an entity can be any person, group, organization, place, or object associated with another entity. While the embodiments in this application are described using individuals associated with one or more social networks as exemplary entities, it is contemplated that these embodiments can be applied to other types of entities. In some embodiments, a social network may be formed by entities or users communicating with each other on a dedicated website or other application that enables entities or users to post information, comments, messages, videos, images, etc. In other embodiments, a social network may be formed by entities or customers utilizing a business platform, wherein some or all of the relationships on the business platform are relationships between entities and businesses.

[0019] In some embodiments, the system's processor may be configured to acquire similarity data associated with multiple entities via a communication network. Similarity data is data representing comparisons between two or more entities. For example, similarity data may include data indicating the frequency of communication between entities (i.e., communication frequency), comparisons of how entities present themselves to social networks (i.e., user profile similarity), comparisons of entities' careers (i.e., job similarity), comparisons of the geographical locations of two entities (i.e., proximity similarity), the rate or amount of currency exchange between two entities (i.e., currency exchange similarity), comparisons of services recommended to each entity (i.e., recommendation similarity), and so on.

[0020] Additionally, the processor can be configured to acquire external data associated with a subset of those entities via a communication network. This subset of entities could be existing customers, thus allowing for the collection of relevant external data. For example, external data could include webpage ranking scores associated with activities performed by both the first and second entities. An entity's webpage ranking score might come from a graph-based algorithm called PageRank, which was originally used to rank websites based on their relative importance to the network. Since its invention, PageRank has been used by social networking companies to rank people using social networks to determine their relative importance on a particular social networking platform. Entities may possess external data whose webpage ranking scores are associated with external sources. The disclosed embodiments utilize external data to provide recommendations to other entities connected to entities that have external data.

[0021] The system's processor can be further configured to train a classification model based on external data and similarity data. To train the classification model, the system's processor can be further configured to determine seed links based on external data. Seed links can represent relationships between two or more seeds or entities that are strongly related to an external source. Seed links can be positive or negative. That is, a positive seed link can identify two seeds that are strongly related to each other, while a negative seed link can identify two seeds that are weakly related to each other. Determining seed links based on external data can also include determining a first seed based on external data of a first entity, determining a second seed based on external data of a second entity, and determining the seed link between the first and second seeds based on external data of the first entity, external data of the second entity, and predetermined values.

[0022] In some embodiments, after seed links are determined based on external data, the system's processor can be configured to determine relationship strength based on similarity data and seed links. Relationship strength can be defined as the strength of the relationship between two seeds relative to other entities. Determining relationship strength based on similarity data and seed links can also include training a classification model using seed links and similarity data to learn relationship weights. The classification model can be stored in memory and trained using supervised or semi-supervised machine learning techniques.

[0023] In some embodiments, the system may be further configured to determine an entity's expected score based on a classification model and to provide recommendations to the entity based on the expected score. For example, this entity may be a potential customer, and therefore, no associated external data exists. Since the entity lacks external data, the expected score may be a prediction of the likelihood that the entity might use or consider recommendations used or considered by seed entities (e.g., existing customers). In some embodiments, external data is associated with external sources that are linked to the provided recommendations. Recommendations may be advertisements, invitations, promotions or discounts on business or social networks, offers to purchase items, etc. Furthermore, the entity's expected score may also be compared with the expected scores of other entities. In some embodiments, the processor may be configured to provide recommendations to one or more entities with the highest expected scores among all entities.

[0024] Reference will now be made in detail to exemplary embodiments, examples of which are shown in the accompanying drawings and disclosed herein. Where convenient, the same reference numerals will be used throughout the drawings to refer to the same or similar parts.

[0025] Figure 1 This is a schematic diagram of an exemplary system environment 100 that provides recommendations based on seed-supervised learning, according to some embodiments of this application.

[0026] refer to Figure 1 The system environment 100 may include various components, such as a communication network 102, a system 105, a social network 108, an external source 110, a device terminal 112, a database 114, a server cluster 116, and a cloud service 118. These different components can be implemented using a variety of different devices, such as supercomputers, servers, personal computers, smartphones, and tablets. Furthermore, these components may include hardware, software, and / or firmware modules.

[0027] like Figure 1As shown, communication network 102 may include one or more interconnected wired or wireless data networks that receive data from one service or device (e.g., system 105) and transmit it to another service or device (e.g., social network 108, external source 110, device terminal 112, database 114, server cluster 116, and cloud service 118). For example, communication network 102 may be implemented as one or more of the Internet, wired wide area network (WAN), wired local area network (LAN), wireless LAN (e.g., IEEE 802.11, Bluetooth, etc.), wireless WAN (e.g., WiMAX), etc. Each component in system environment 100 may communicate bidirectionally with other components of system environment 100 via communication network 102 or via one or more direct communication links (not all are shown).

[0028] System 105 can be configured to provide recommendations based on seed-supervised learning. In some embodiments, system 105 may obtain data from social network 108 or external source 110. Furthermore, in some embodiments, data from social network 108 or external source 110 may be stored on system 105. In some embodiments, social network 108 may be formed by entities or users communicating with each other on a dedicated website or other application that enables entities or users to post information, comments, messages, videos, images, etc. In other embodiments, social network 108 may be formed by entities or customers utilizing a business platform, where some or all of the relationships on the business platform are relationships between entities and businesses. As noted in this application, an entity can be any person, group, organization, place, or object that is associated with another entity. While the embodiments in this application are described using people associated with one or more social networks as exemplary entities, it is contemplated that these embodiments can be applied to other types of entities. As described above, social network 108 may provide data to system 105. This data may include similarity data associated with each entity of social network 108. Similarity data is data representing comparisons between two or more entities.

[0029] External source 110 can be a company, enterprise, or website that facilitates another social network. In some embodiments, external source 110 may be associated with the same social network 108; in others, it may be different. In any case, like social network 108, external source 110 can provide data to system 105. This data may include external data of a subset of entities in social network 108.

[0030] Device terminal 112 may be a supercomputer, server, personal computer, or mobile device such as a smartphone or tablet. Device terminal 112 may be configured to receive input from a user for transmission to system 105. Device terminal 112 may also receive input from system 105. For example, system 105 may enable device terminal 112 to alert the user by sending a notification or any other means that can be determined by those skilled in the art.

[0031] Database 114 may be configured to store information consistent with the disclosed embodiments. In some aspects, components of system environment 100 (shown and not shown) may be configured to receive, acquire, collect, gather, generate, or produce information to be stored in database 114. For example, in some embodiments, components of system environment 100 may receive or acquire information for storage via communication network 102. For example, database 114 may store data associated with one or more entities. In other aspects, components of system environment 100 may store information in database 114 without using communication network 102 (e.g., via a direct connection). In some embodiments, components of system environment 100 (including, but not limited to, system 105) may use information stored in database 114 for processes consistent with the disclosed embodiments.

[0032] Server clusters 116 can be located in the same data center or in different physical locations. Multiple server clusters 116 can be formed as a mesh to share resources and workloads. Each server cluster 116 may include multiple linked nodes that collaborate to run various applications, software modules, analytics modules, rule engines, etc. Each node can be implemented using various different devices, such as supercomputers, personal computers, servers, mainframes, mobile devices, etc. In some embodiments, the number of servers and / or server clusters 116 can be increased or decreased based on workload.

[0033] Cloud service 118 may include physical and / or virtual storage systems associated with cloud storage for storing data and providing access to the data via public networks such as the Internet. Cloud service 118 may include cloud services provided by, for example, Amazon, Apple, Cisco, Citrix, IBM, Joyent, Google, Microsoft, Rackspace, Salesforce.com, Verizon / Terremark, etc., or other types of cloud services accessible via communication network 102. In some embodiments, cloud service 118 includes multiple computer systems spanning multiple locations and having multiple databases or multiple geographic locations associated with a single or multiple cloud storage services. As used herein, cloud service 118 refers to the physical and virtual infrastructure associated with a single cloud storage service. In some embodiments, cloud service 118 manages and / or stores data associated with providing recommendations to entities.

[0034] Figure 2 This is a block diagram of an exemplary system 105 that provides recommendations based on seed-supervised learning, according to some embodiments of this application. Figure 2 As shown, system 105 may include memory 200, processor 210, and communication interface 250. Processor 210 may also include multiple modules, such as a relation learning unit 212, a recommendation unit 214, and a testing unit 216. These modules (and any corresponding submodules or subunits) may be functional hardware units of processor 210 (e.g., portions of integrated circuits) designed to be used with other components or as part of a program (stored on a computer-readable medium) that performs one or more functions by processor 210. Although Figure 2 Units 212-216 shown are all within a single processor 210, but it is conceivable that these units could be distributed across multiple processors, either close to or far from each other. System 105 can be implemented in a cloud service 118, device 112, or a single computer / server (e.g., server cluster 116). In some embodiments, one or more of these modules can be combined.

[0035] System 105 can provide recommendations in two phases. In the first phase, also known as the “training phase,” system 100 can use data associated with social network 108 and external source 110 to generate and train a classification model; and in the second phase, also known as the “recommendation phase,” system 100 can apply the trained relational model to provide recommendations for entities.

[0036] Communication interface 250 can establish one or more sessions with social network 108, external source 110, and device terminal 112 via communication network 102. In some embodiments, communication interface 250 can continuously receive data from social network 108, external source 110, and device terminal 112 via communication network 102. Furthermore, communication interface 250 can transmit data between any of social network 108, external source 110, and device terminal 112 and any of relationship learning unit 212, recommendation unit 214, and testing unit 216 via communication network 102.

[0037] The relation learning unit 212 can perform a portion of the training phase, meaning it can use data associated with the social network 108 and external sources 110 to generate and train a classification model. In some embodiments, the relation learning unit 212 can acquire similarity data associated with multiple entities via the communication network 102. Additionally, the relation learning unit 212 can acquire external data associated with a subset of entities via the communication network 102. In some embodiments, the relation learning unit 212 can train a classification model based on the acquired entity information, including external data and similarity data.

[0038] Recommendation unit 214 may perform a portion of the recommendation phase by using a classification model trained in the training phase and similarity data of one or more entities to provide recommendations to one or more entities. For example, recommendation unit 214 may determine the expected score of one or more entities based on the classification model. Recommendation unit 214 may also provide recommendations to one or more entities based on the expected score via communication network 102. In some embodiments, one or more entities may not have any external data.

[0039] Test unit 214 can monitor or test classification models, such as those trained by relation learning unit 212. For example, test unit 214 can test to see if the model converges mathematically. If it does not converge mathematically, the classification model can cause system 105 to send a notification about the problem to device terminal 112 via communication network 102. A classification model that does not converge mathematically cannot be used; therefore, test unit 214 can cause system 105 to ignore the previously trained model and train a new classification model. Additionally, in some embodiments, test unit 214 can monitor or test the classification model while relation learning unit 212 is training it. However, in other embodiments, test unit 214 can only monitor or test the classification model after relation learning unit 212 has finished training it.

[0040] Figure 3 This is a diagram of an exemplary social network 108 shown according to some embodiments of this application. For example... Figure 3As shown, social network 108 may include multiple entities, such as 302a-302f. As indicated in this application, an entity can be any person, group, organization, place, or object associated with another entity. While the embodiments in this application are described using individuals associated with one or more social networks as exemplary entities, it is contemplated that these embodiments can be applied to other types of entities. In some embodiments, social network 108 may be formed by entities or users communicating with each other on a dedicated website or other application that enables entities or users to post information, comments, messages, videos, images, etc. In other embodiments, social network 108 may be formed by entities or customers utilizing a business platform, wherein some or all of the relationships on the business platform are relationships between entities and businesses.

[0041] Each entity 302a-302f may also have relations 304ab-304df with another entity. Consistent with this application, relation 304ij represents the relationship between entities 302i and 302j, where i and j = a, b, c, d, e, f, g, and f. The relationship between two entities in one group (e.g., 302a and 302b) may be weaker or stronger than the relationship between two entities in another group (e.g., 302b and 302e). This relative measure of the relationship between two entities can be referred to as "relationship strength".

[0042] In some cases, as shown in the figure, it may be helpful to use graph theory to describe social network 108. In graph theory, there exists a set of vertices (e.g., entities 302a-302f) connected by a set of edges (relations 304ab-304df), which can be represented by the equation G = (V, E), where social network 108 or its equivalent graph (G) comprises multiple entities 302a-302f or vertices (V) and one or more relations 304ab-304df or edges (E). Furthermore, each relation or edge within social network 108 may have a relation strength or edge weight (Wij), where i and j represent the specific entities connected by the relation.

[0043] Figure 4 This is a flowchart of an exemplary process 400 for providing recommendations based on seed-supervised learning. For example, process 400 may be implemented by system 105, specifically by processor 210. Process 400 may include steps 410-450 as described below. In some embodiments, relation learning unit 212 may perform steps 410-430, and recommendation unit 214 may perform steps 440-450.

[0044] In step 410, system 105 may acquire similarity data associated with multiple entities. The similarity data may be data indicating or corresponding to comparisons between two or more entities. Such comparisons may be represented numerically. In some embodiments, similarity may include various data indicating or corresponding to different types of comparisons. For example, similarity data may include data indicating communication frequency, similarity of user profiles, job similarity, proximity similarity, currency exchange similarity, recommendation similarity, etc.

[0045] Communication frequency can indicate the frequency or rate of communication between two entities. In some embodiments, communication frequency can indicate the number of times entities communicate within a set time period. For example, communication frequency can indicate that entities communicate once a month, twice a week, forty times a year, etc. Communication frequency can also include various forms of communication, including calls, text messages, instant messaging, emails, etc. In addition, communication frequency can include other forms of communication, such as posting information, comments, messages, videos, images, etc. The two entities can have communication frequencies for each form of communication. In some embodiments, communication frequency can also include a weighted and combined communication frequency of various communication forms.

[0046] User profile similarity can indicate how entities present themselves to comparisons on social networks. For example, user profile similarity can include a comparison of how two entities change their user profiles on each entity's social network over time. As another example, user profile similarity can include information from user profiles provided to the network by two entities (e.g., name, address, phone number, education, occupation, interests, etc.). Furthermore, user profile similarity can also include a comparison of the different interactions each entity makes on social network 108. This can include social media activity such as liking pages, commenting on various posts, sharing posts, using emojis, and so on.

[0047] Furthermore, job similarity can indicate comparisons of entities' careers based on information provided to the social network 108. For example, job similarity can include comparisons of salary, bonuses, taxes, job roles, job functions, professional networks, etc.

[0048] Furthermore, proximity similarity can indicate a comparison of the geographic locations of two entities provided to the social network 108. Proximity similarity can include a comparison of geographic locations at any time, place, manner, or any combination thereof. The comparison can be the distance between two entities, the distance between two entities and one or more reference points, the velocity difference between two entities, etc. Proximity similarity can also be captured using hardware or software, such as GPS and cellular tracking in devices owned by the entities.

[0049] Furthermore, currency exchange similarity can indicate the ratio or amount of currency exchanged between two entities. Currency exchange similarity can include one or more types of currency. For example, currency exchange can be one or more internet currencies. In some embodiments, the currency is a real currency in various forms as defined by a specific country (e.g., RMB, USD, GBP, etc.). Currency exchange similarity can also include the ratio or amount of currency exchanged through red packet sharing (distributing virtual red packets between two entities via the internet). Further, currency exchange can refer to different categories of currency exchange between two entities, such as gifts, payments, loans, etc.

[0050] Recommendation similarity can indicate a comparison between services recommended by two entities. These services can be recommended by entities to each other or to other entities on the social network 108. Recommendation similarity may represent a comparison between services recommended by two entities by category (e.g., non-profit / for-profit, technology, business, finance, products offered, etc.). Recommendation similarity can also indicate whether two entities have viewed, skipped, commented on, liked, or shared the same one or more recommendations.

[0051] In some embodiments, system 105 may obtain similarity data directly from social network 108. In other embodiments, system 105 may obtain similarity data directly from a third-party source that has collected data associated with social network 108. Further, in other embodiments, system 105 may obtain similarity data from database 114, server 116, or cloud service 118.

[0052] In step 420, system 105 may acquire external data associated with a subset of entities. The acquired external data may include data associated with external source 110 and / or data associated with third-party sources. The external data may include the page ranking score of each entity in the subset of entities associated with the activity. The activity performed may include making a purchase, selling an item, joining a service, watching an advertisement, meeting at a location, etc. This activity may be performed by entities interacting with businesses, software platforms, devices, etc. Interacting with a business may include, for example, watching advertisements associated with the business, purchasing products from the business, liking, sharing, watching, or commenting on posts associated with the business, opening emails from the business, etc. Furthermore, the entity's page ranking score may come from a graph-based algorithm called PageRank, which was originally used to rank websites based on their relative importance to the network. Here, each entity's page ranking score can indicate how an entity ranks relative to all entities that have external data associated with the performed activity.

[0053] In some embodiments, system 105 may obtain external data directly from external source 110. In other embodiments, system 105 may obtain external data directly from another source that has already collected data associated with external source 110. Further, in other embodiments, system 105 may obtain external data from database 114, server 116, or cloud service 118.

[0054] In step 430, system 105 can train a classification model based on external data and similarity data. System 105 can use semi-supervised learning to train the classification model. Semi-supervised learning involves using a small number of known features from a small subset of data (e.g., external data of entities) to predict features from a large amount of data (entities).

[0055] Here, using graph theory terminology may be helpful. For example, as mentioned above, a social network 108 (G) includes multiple entities 302a-302f (V) and one or more relations 304ab-304df (E), and each relation has a relation strength denoted by (Wij), where i and j represent specific entities connected by the relation. It should be understood that the disclosed similarity data can also be identified using graph theory terminology. For example, since similarity data can be data indicating comparisons between two or more entities, similarity data can be the determination of the characteristics of relations between entities (V) or edges (E). Thus, similarity data or edge characteristics can be represented as fe,i, where e represents relations 304ab-304df or edges, and i represents a specific type of similarity data (e.g., communication frequency, etc.).

[0056] Back Figure 5 It provides a method for implementing the embodiments consistent with those disclosed. Figure 4 The flowchart illustrates an exemplary process for step 430. In step 510, system 105 may determine one or more seed links based on external data. A seed link may represent a relationship between two or more seeds. A seed is an entity that has a close relationship with an external source. A seed link may have a property (i.e., a label also called a seed link or edge), where the property of the seed link can be positive or negative. That is, a positive seed link can identify two seeds that have a strong relationship with each other, while a negative seed link can identify two seeds that have a weak relationship.

[0057] Reference Figure 6According to some embodiments of this application, flowcharts are provided for an exemplary process for implementing step 510. For example, at 610, system 105 may determine multiple seeds from a subset of entities based on external data. Determining multiple seeds from a subset of entities may involve system 105 comparing the external data for each entity with a threshold. For example, system 105 may compare the webpage ranking score of an entity associated with an activity with a webpage ranking score threshold. In some embodiments, if an entity's webpage ranking score exceeds or is at least equal to the webpage ranking score threshold, the entity is determined to be a seed. In these embodiments, if an entity's webpage ranking score fails to exceed or is at least equal to the webpage ranking score threshold, the entity is not determined to be a seed. System 105 may also apply other techniques involving multiple webpage ranking scores. For example, system 105 may calculate a weighted average of the webpage ranking scores and compare the weighted average with a threshold to determine a seed.

[0058] At 620, system 105 can iteratively determine whether a seed is related to another seed. For example, system 105 can compare all combinations of seeds to find whether the first seed has a subset of similarity data that overlaps with the similarity data of the second seed. If a seed shares a subset of overlapping similarity data with another seed, it is determined that the seed is related to the other seed. As another example, additional data such as the programming structure (link list, hash map, array, etc.) used to store relationships between entities in social network 108 can be stored in memory 200. System 105 can search the programming structure to determine whether a relationship exists between the first seed and the second seed. If a seed is not related to another seed, there is no seed link (step 630), and system 105 can proceed to the next seed (step 660).

[0059] However, if a seed is indeed related to another seed, system 105 can determine the characteristics of the link between the seeds (or seed link characteristics) (step 640). For example, the characteristics of a seed link can be positive or negative. For example, system 105 can use external data and predetermined values ​​to determine the characteristics of a seed link. In some embodiments, system 105 can determine the proximity of external data linking the seeds. For example, system 105 can calculate the relative difference between the linked seeds and determine the characteristics of the seed link. For example, if the relative difference exceeds a predetermined value (such as a proximity tolerance threshold), system 105 can determine the characteristics of the seed link as positive. However, if the relative difference does not exceed the predetermined value, system 105 can determine the characteristics of the seed link as negative. In other examples, the opposite can also be true, where system 105 can determine the characteristics of a seed link as positive when the relative difference does not exceed the predetermined value, or as negative when the relative difference exceeds the predetermined value.

[0060] In other embodiments, system 105 can determine the characteristics of a seed link. For example, if the maximum or minimum score between two seeds exceeds a predetermined value, such as an indifference threshold, system 105 can determine the characteristic of the seed link as positive. However, if the maximum or minimum score between two seeds does not exceed the predetermined value, system 105 can determine the characteristic of the seed link as negative. In other examples, the opposite may also be true: system 105 can determine the characteristic as positive when the maximum or minimum score between two seeds does not exceed the predetermined value, or determine the characteristic as negative when the maximum or minimum score between two seeds exceeds the predetermined value.

[0061] Once system 105 has determined the characteristics (or seed link properties) of a seed link, it can store the seed link (650) as positive or negative for use in subsequent step 520. After storing each seed link with its respective characteristics, system 105 can proceed to the next seed (step 660) if a next seed exists.

[0062] Back Figure 5 In step 520, system 105 can determine the relationship strength between entities 302a-302f in social network 108 based on similarity data and seed links. System 105 can use logistic regression to train a classification model to learn the relationship strength or edge weight (We) of multiple entities 304a-304f based on similarity data (fe,i) and seed link characteristics (ye) (i.e., labels). System 105 can train a classification model to calculate one or more relationship strengths using the following mathematical representation.

[0063]

[0064] Then, the sigmoid function in the linear model can be applied to give the predicted relation strength for any relation (E) 304ab-304df that has similar data (fe,i) but does not have a corresponding seed link.

[0065]

[0066] Therefore, the predicted relation strength of a relation might be the probability of belonging to two seeds (i.e., entities or vertices) in external data. Returning to... Figure 4 As shown, in step 440, system 105 can determine the expected score of one or more entities based on a classification model. For example, system 105 can use the inbound normalized PageRank function to calculate the expected score. This function can be represented as follows:

[0067] The normalization function is:

[0068] It should be understood that the denominator of the normalization function has a maximum value of 1, which ensures that when the strength of the inbound relation to the entity is small, the system 105 will not artificially amplify the strength of the weak relation.

[0069] System 105 can also add the transmission probability vector (d) and reset strength or relation (r) to the inbound normalized PageRank function, as described below.

[0070]

[0071] .

[0072] The transmission probability vector (d) and relation strength reset (r) allow system 105 to consider the inherent randomness in the inbound normalized PageRank function when needed.

[0073] Further, in step 450, system 105 may provide recommendations based on the expected scores of one or more entities. Recommendations may be advertisements, invitations, promotions or discounts, suggestions to purchase something, etc. Recommendations may be provided by businesses, social networks, entities, etc. Recommendations may be for an activity. Similarly, the activity may include making a purchase, selling an item, joining a service, watching an advertisement, meeting at a location, etc. The recommendation may be for an external source 110 or for a product of external source 110. In some embodiments, system 105 has already obtained external data associated with the recommendation from external source 110. System 105 may provide recommendations to entities in various ways. System 105 may provide recommendations to entities with the highest expected scores, where high expected score indicates an expectation that the target entity will view or participate in the recommendation. System 105 may provide recommendations over a period of time. System 105 may also repeat steps 410-440 to provide updated expected scores at predetermined intervals or in real time. Further, if the similarity data or external data has changed, system 105 may also repeat steps 410-440. It should be understood that expected scores can improve the accuracy of the recommendations provided.

[0074] The description of the disclosed embodiments is not exhaustive and is not limited to the precise forms or embodiments disclosed. Modifications and adaptations to the embodiments will be apparent from consideration of the specification and practice of the disclosed embodiments. For example, the described implementations include hardware, firmware, and software, but systems and techniques consistent with this application can be implemented solely as hardware. Furthermore, the disclosed embodiments are not limited to the examples discussed herein.

[0075] Computer programs based on the written descriptions and methods of this specification fall within the skill scope of software developers. Various programming techniques can be used to create a wide variety of programs or program modules. For example, program segments or modules can be designed or used in Java, C, C++, assembly language, or any such programming language. One or more such software parts or modules can be integrated into a computer system, a non-transitory computer-readable medium, or existing communication software.

[0076] Furthermore, although illustrative embodiments are described herein, the scope includes any and all embodiments based on this application that have equivalent elements, modifications, omissions, combinations (e.g., aspects of various embodiments), adaptations, or alterations. Elements in the claims will be interpreted broadly based on the language used in the claims, and not limited to the examples described in this specification or in the course of the application, which will be interpreted as non-exclusive. Moreover, the steps of the disclosed methods can be modified in any way, including by reordering steps or inserting or deleting steps. Therefore, the specification and embodiments are intended to be considered exemplary only, and the true scope and spirit are indicated by the full scope of the appended claims and their equivalents.

Claims

1. A system for providing recommendations to entities, comprising: processor; and A memory device for storing instructions, which, when executed by the processor, cause the processor to: Similarity data associated with a first entity, a second entity, and a third entity is obtained through a communication network, wherein the similarity data indicates comparisons between different entities; External data associated with the first entity and the second entity is acquired through the communication network; A classification model is trained based on the external data and the similarity data; The expected score of the third entity is determined based on the classification model. and Recommendations based on the expected scores are provided to the third entity through the communication network; The external data includes webpage ranking scores associated with activities performed by both the first entity and the second entity.

2. The system according to claim 1, characterized in that, The first, second, and third entities are associated with a social network.

3. The system according to claim 1, characterized in that, The similarity data includes one or more data indicating the following: communication frequency, similarity of user profiles, similarity of jobs, similarity of lifestyles, similarity of proximity, similarity of currency exchange, and similarity of recommendation services.

4. The system according to claim 1, characterized in that, The processor is also configured to: The characteristics of the seed link are determined based on the external data, wherein the characteristics of the seed link are positive or negative; The strength of the relationship is determined based on the similarity data and the characteristics of the seed link.

5. The system according to claim 4, characterized in that, The processor is also configured to: The first seed is determined based on the external data of the first entity; The second seed is determined based on the external data of the second entity; and Based on the external data of the first entity, the external data of the second entity, and a predetermined value, determine whether the characteristic of the seed link between the first and second seeds is positive or negative.

6. The system according to claim 4, characterized in that, The determined relationship strength is the relationship strength between the first entity and the third entity.

7. The system according to claim 1, characterized in that, The processor is further configured as follows: The expected score of the fourth entity is determined based on the classification model, wherein the similarity data is further associated with the fourth entity, and The expected score of the third entity tends to the expected score of the fourth entity.

8. The system according to claim 7, characterized in that, The processor is also configured to provide recommendations to the fourth entity based on the expected score of the third entity and the expected score of the fourth entity.

9. The system according to claim 1, characterized in that, The recommendation is directed to an external source, and the external source provides the external data.

10. A computer-implemented method for providing recommendations to entities, comprising: Similarity data associated with a first entity, a second entity, and a third entity is obtained through a communication network, wherein the similarity data indicates comparisons between different entities; External data associated with the first entity and the second entity is acquired through the communication network; A classification model is trained based on the external data and the similarity data; The expected score of the third entity is determined based on the classification model. and The recommendation based on the expected score is provided to the third entity via the communication network; the external data includes webpage ranking scores associated with activities performed by both the first entity and the second entity.

11. The method according to claim 10, characterized in that, The first, second, and third entities are associated with a social network.

12. The method according to claim 10, characterized in that, The similarity data includes one or more data indicating the following: communication frequency, similarity of user profiles, similarity of jobs, similarity of lifestyles, similarity of proximity, similarity of currency exchange, and similarity of recommendation services.

13. The method according to claim 10, characterized in that, Training a classification model based on the similarity data and the external data also includes: The characteristics of the seed link are determined based on the external data, wherein the characteristics of the seed link are positive or negative; The strength of the relationship is determined based on the similarity data and the characteristics of the seed link.

14. The method according to claim 13, characterized in that, The characteristics of determining seed links further include: The first seed is determined based on the external data of the first entity; The second seed is determined based on the external data of the second entity; and Based on the external data of the first entity, the external data of the second entity, and a predetermined value, determine whether the characteristic of the seed link between the first and second seeds is positive or negative.

15. The method according to claim 13, characterized in that, The determined relationship strength is the relationship strength between the first entity and the third entity.

16. The method according to claim 10, characterized in that, Also includes: The expected score of the fourth entity is determined based on the classification model, wherein the similarity data is further associated with the fourth entity, and The expected score of the third entity tends to the expected score of the fourth entity.

17. The method according to claim 16, characterized in that, It further includes providing recommendations to the fourth entity based on the expected score of the third entity and the expected score of the fourth entity.

18. A non-transitory computer-readable medium storing a set of instructions, which, when executed by at least one processor of a recommendation system, cause the recommendation system to perform a method for providing recommendations to entities, the method comprising: Obtain similarity data associated with a first entity, a second entity, and a third entity, wherein the similarity data indicates comparisons between different entities; Obtain external data associated with the first entity and the second entity; A classification model is trained based on the external data and the similarity data; The expected score of the third entity is determined based on the classification model. and Recommendations are provided to the third entity based on the expected score; the external data includes webpage ranking scores associated with activities performed by both the first and second entities.

Citation Information

Patent Citations

  • Information recommending method and information recommending device in social media

    CN104281622A