Distributed SLB service recommendation method and system based on incentive mechanism
By introducing an incentive mechanism in the distributed SLB recommendation system, determining the incentive cost and collaboration proportional relationship, identifying the collection of client participants, and issuing incentives based on marginal costs, the problems of insufficient motivation for client participation and privacy leakage in distributed learning are solved, and effective distributed SLB service recommendation and privacy protection are achieved.
Patent Information
- Application Number
- CN202411755987.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-03
AI Technical Summary
The existing distributed SLB algorithm ignores the costs faced by terminal devices participating in distributed learning in the recommendation system, lacks incentive mechanism design, resulting in insufficient motivation for client participation and risk of privacy leakage.
A distributed SLB service recommendation method based on incentive mechanism is proposed. By obtaining client data, a distributed SLB model is constructed, the incentive cost and collaboration proportional relationship is determined, the client participants collection is identified, and incentives are issued according to marginal costs, and the client uploads data is obtained to achieve the recommended results.
The incentive mechanism encourages clients to participate, meet the privacy needs of clients before obtaining incentives, and resists malicious users to obtain profits by falsely reporting costs, achieving the effectiveness and privacy protection of distributed SLB service recommendations.
Smart Images

Figure CN119939010A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of distributed personalized recommendation technology, and in particular to a distributed SLB service recommendation method and system based on an incentive mechanism. Background Art
[0002] In recent years, various media applications on the Internet have provided more and more rich media content for users to browse. One of the most important tasks of these applications is to recommend content that meets user preferences in order to attract and retain users. For example, popular video platforms such as TikTok provide timely and accurate content recommendations by inferring users' personal information, browsing history, and other data. Similarly, online learning platforms such as Coursera collect users' course selections, learning progress, grades, etc., and use machine learning algorithms to recommend appropriate courses and learning content.
[0003] With the continuous improvement of terminal device computing power, the paradigm of collaborative learning using multiple clients has been widely used, which is called distributed learning technology. However, distributed learning requires multiple clients to transmit local data or models. Some malicious attackers can directly obtain or indirectly infer the private information of client users through these data or models, such as gender, age, preferences, etc. Therefore, clients need to consider the privacy leakage costs brought by participating in distributed learning. At the same time, the participation of clients in distributed learning will also incur local learning costs and communication costs. In real scenarios, terminal devices need to consider their own benefits. The motivation of clients to participate can only be rationalized if they receive sufficient incentives. Therefore, distributed learning algorithms need to consider the design of incentive mechanisms.
[0004] At present, related methods extend the SLB algorithm to distributed recommendation systems. These distributed methods require the client to upload data or models related to actions and rewards, and then the server aggregates and synchronizes the models. However, most of these methods assume that the client participates voluntarily, ignoring the various costs faced by terminal devices participating in distributed learning, and lack the design of incentive mechanisms to rationalize the client's participation motivation. A few studies have introduced incentive mechanisms in distributed SLB algorithms, but these mechanisms also require the client to upload relevant data about the action in advance to formulate appropriate incentives. This part of the data still has privacy concerns.
[0005] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the invention
[0006] The main purpose of the embodiments of the present application is to propose a distributed SLB service recommendation method and system based on an incentive mechanism, which can motivate the participation of clients through the incentive mechanism, meet the privacy needs of clients before obtaining incentives, and resist malicious users from defrauding higher profits by falsely reporting costs.
[0007] To achieve the above purpose, one aspect of an embodiment of the present application proposes a distributed SLB service recommendation method based on an incentive mechanism, the method comprising:
[0008] Obtain client data and build a distributed SLB model, solve it using the ridge regression estimation algorithm, and determine the amount of data uploaded by the client;
[0009] Determine the incentive cost and construct a collaborative proportion relationship according to the amount of data uploaded by the client, and determine the client participant set;
[0010] The marginal cost is determined, and the server issues incentives to the clients in the client participant set according to the marginal cost, obtains the uploaded data of the clients, and obtains the distributed SLB service recommendation result.
[0011] In some embodiments, obtaining client data and building a distributed SLB model, solving the problem using a ridge regression estimation algorithm, and determining the amount of data uploaded by the client include:
[0012] Acquire client data, and determine historical action information of the client data and historical reward information of the client data;
[0013] Constructing a covariance matrix according to the historical action information of the client data and constructing a reward vector according to the historical action information and historical reward information of the client data;
[0014] Constructing a distributed SLB model according to the covariance matrix and the reward vector, and estimating it through a ridge regression estimation algorithm to obtain an estimation result of the distributed SLB model;
[0015] Based on the optimistic strategy, the action with the highest reward confidence upper bound of the estimated result of the distributed SLB model is selected, and the corresponding reward is obtained as the uploaded data of the client, and the amount of uploaded data of the client is recorded.
[0016] In some embodiments, the client data includes first type data, second type data and third type data, wherein:
[0017] The first type of data represents local data newly added by the client after the last data service communication. If the newly added local data has incentives, it is aggregated into the third type of data. If the newly added local data does not have incentives, it is aggregated into the second type of data.
[0018] The second type of data represents local data that has not been uploaded to participate in data service communications;
[0019] The third type of data represents synchronization data used for participating in data service communications.
[0020] In some embodiments, determining the incentive cost and constructing a collaboration ratio relationship according to the amount of data uploaded by the client to determine the client participant set includes:
[0021] Constructing a cooperation ratio relationship according to the amount of data uploaded by the client;
[0022] Determine the incentive costs of the clients and sort them in ascending order to obtain sorted incentive costs of the clients;
[0023] A threshold ratio is set, and the ranked client incentive costs are used to sequentially divide the clients that satisfy the cooperation ratio relationship into the client participant set.
[0024] In some embodiments, the expression of the communication trigger condition is specifically as follows:
[0025]
[0026] In the above formula, Δt k Indicates the time from the last communication to the current round of communication, represents the discrimination threshold, represents the covariance matrix containing the uploaded data, Represents the covariance matrix of the first type of data.
[0027] In some embodiments, the expression of the cooperation ratio relationship is specifically as follows:
[0028]
[0029] In the above formula, α represents the cooperation ratio parameter, P k,t represents the cumulative number of samples participating in the collaboration, Q k,t Represents the total number of samples in the current system, α h Represents the threshold ratio.
[0030] In some embodiments, the determining of the marginal cost, the server issuing incentives to the clients in the client participant set according to the marginal cost, obtaining the uploaded data of the clients, and obtaining the distributed SLB service recommendation result include:
[0031] Determining marginal costs according to the custom reported costs of the client;
[0032] According to the real cost of data and the custom reporting cost, the clients are divided into malicious clients and honest clients;
[0033] Setting a communication trigger condition according to a covariance matrix of the amount of data uploaded by the client;
[0034] Based on the communication triggering condition, issuing incentives to the malicious client and the honest client according to the marginal cost;
[0035] If the malicious client and the honest client are selected to communicate, the covariance matrix and the reward vector corresponding to the newly added local data are uploaded to the server for data communication aggregation to obtain a first service recommendation result;
[0036] If the malicious client and the honest client are not selected to communicate, local data aggregation is performed on the covariance matrix corresponding to the newly added local data and the reward vector to obtain a second service recommendation result;
[0037] The distributed SLB service recommendation result is obtained by combining the first service recommendation result and the second service recommendation result.
[0038] In some embodiments, the marginal cost is calculated by keeping the custom reporting cost of the remaining clients unchanged, assuming that the custom reporting cost of the selected client changes, and changing it in turn to the custom reporting cost of each unselected client, and repeatedly determining the set of client participants so that the minimum custom reporting cost at which the selected client will not be selected by the server is the marginal cost.
[0039] In some embodiments, the issuing of incentives to the malicious client and the honest client according to the marginal cost includes:
[0040] For the honest client, if the custom reporting cost is less than the marginal cost, the honest client is selected to perform data service communication and obtain the incentive; if the custom reporting cost is greater than the marginal cost, the honest client is not selected to perform data service communication and cannot obtain the incentive;
[0041] For the malicious client, if the custom reporting cost is less than the real cost of the data and less than the marginal cost, or the real cost of the data is less than the custom reporting cost and less than the marginal cost, or the custom reporting cost is less than the marginal cost and less than or equal to the real cost of the data, then the malicious client is selected for data service communication and obtains the incentive; if the real cost of the data is less than the marginal cost and less than or equal to the custom reporting cost, or the marginal cost is less than the custom reporting cost and less than or equal to the real cost of the data, or the marginal cost is less than the real cost of the data and less than or equal to the custom reporting cost, then the malicious client is not selected for data service communication and cannot obtain the incentive.
[0042] To achieve the above purpose, another aspect of the embodiment of the present application proposes a distributed SLB service recommendation system based on an incentive mechanism, the system comprising:
[0043] The first module is used to obtain client data and build a distributed SLB model, and solve it through the ridge regression estimation algorithm to determine the amount of data uploaded by the client;
[0044] The second module is used to determine the incentive cost and construct a collaboration ratio relationship according to the amount of data uploaded by the client to determine the client participant set;
[0045] The third module is used to determine the marginal cost. The server issues incentives to the clients in the client participant set according to the marginal cost, obtains the uploaded data of the clients, and obtains the distributed SLB service recommendation result.
[0046] The embodiments of the present application include at least the following beneficial effects: The present application provides a distributed SLB service recommendation method and system based on an incentive mechanism. The scheme obtains client data and constructs a distributed SLB model, solves it through a ridge regression estimation algorithm, determines the client's uploaded data and the amount of uploaded data, and then determines the incentive cost and constructs a collaborative proportional relationship based on the client's uploaded data amount, determines the set of client participants, and takes into account the model heterogeneity of distributed learning. An incentive mechanism is designed to incentivize client participation and meet the client's privacy needs before obtaining incentives. The marginal cost is further determined. The server issues incentives to the client based on the marginal cost, obtains the client's uploaded data, and is used to find the optimal recommendation strategy. At the same time, the client's privacy protection needs for data such as actions and rewards are guaranteed until the server provides satisfactory incentives. The incentive mechanism is designed to meet individual rationality and incentive compatibility, which can resist malicious users from defrauding higher returns by falsely reporting costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a flow chart of a distributed SLB service recommendation method based on an incentive mechanism provided in an embodiment of the present application;
[0048] Figure 2 It is a structural diagram of a distributed SLB service recommendation system based on an incentive mechanism provided in an embodiment of the present application;
[0049] Figure 3 It is a schematic diagram of a distributed SLB service recommendation model constructed in an embodiment of the present application;
[0050] Figure 4 It is a schematic diagram of a distributed personalized recommendation algorithm provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of systems and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.
[0052] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".
[0053] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0055] Before describing the embodiments of the present application in detail, some nouns and terms involved in the embodiments of the present application are first described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0056] 1) Model heterogeneity: The differences in local models between clients are called model heterogeneity. This will cause the overall model obtained through collaboration to deviate significantly from the local model, causing the decisions generated by the recommendation algorithm to deviate from the optimal decision, affecting the performance of the recommendation system. This is also an important issue that distributed algorithms need to consider.
[0057] 2) Stochastic Linear Bandit (SLB): A common method for personalized recommendation is to construct it as a stochastic linear bandit problem, in which candidate media content is constructed as a set of candidate actions, and the user feedback obtained by recommending a certain media content is constructed as a reward. The linear model parameters are learned using historical actions and rewards, and the optimal decision for the next round is predicted based on the model parameters.
[0058] There are some deficiencies in the related technologies, such as the existing models ignore the various costs faced by terminal devices participating in distributed learning, lack the design of incentive mechanisms to rationalize the client's motivation to participate, and there are problems with privacy leakage.
[0059] In view of this, a distributed SLB service recommendation method based on an incentive mechanism is provided in the embodiment of the present application. The incentive mechanism is introduced in distributed learning, and the preference differences of recommendation system users in different regions are characterized through a multi-task structure, and the privacy requirements of some institutions that do not want data to be disclosed are met. Based on this model, a distributed personalized recommendation algorithm is proposed to find the optimal recommendation strategy, while ensuring the privacy protection requirements of the client for data such as actions and rewards until the server provides satisfactory incentives. In addition, the algorithm designs an incentive mechanism to meet individual rationality and incentive compatibility, which can resist malicious users from defrauding higher returns by falsely reporting costs.
[0060] First of all, it should be noted that Figure 3 As shown, an embodiment of the present invention proposes a multi-task distributed learning model for a personalized recommendation system. Horizontally, each client group corresponds to a specific region, such as a city, town, or institution. These regions differ in the number of users, participation levels, and preference patterns, and each region is represented as a different task to capture model heterogeneity between regions. Vertically, the framework combines a distributed learning architecture with an integrated incentive mechanism to promote client participation, wherein the system consists of a server and K client groups corresponding to K tasks, where the client set corresponding to task k∈{1,2,…,K} is denoted as The base number is All users are represented as The base number is The global unknown model parameters shared by users in are denoted as θ k , which is the key parameter of the linear reward structure in the SLB model.
[0061] Further, the client model of the embodiment of the present invention is described:
[0062] In real-world scenarios, recommendation behavior on the client does not always occur, that is, the local learning process on the client does not always occur. We assume that recommendation behavior on the client occurs intermittently, rather than every round within a time frame. In each round, only active clients interact with the media files and perform local learning, while inactive clients only maintain previous data in the round without performing further operations. In order to describe this intermittent interaction of the client, we construct the following client model.
[0063] Denote the round limit as T, client group The client index in is represented as For client i k , and denote its state at each time step t∈{1, 2, ..., T} as If client i k If it is active at round t, otherwise The round number of the last communication is represented by t last .use Indicates the number of client i since the last round of communication k The number of active times is defined as Intuitively, It can be regarded as client i k The number of new samples added since the last round of communication, reflecting the client i k The value you can bring by participating in the next round of communications.
[0064] Will Represented as client groups The set of clients active in round t. In round t, for active clients Many media files are candidates for interaction, and these candidates are constructed into action sets. Then the client i k According to the actions and results of historical interactions, select the actions to be performed in this round Then, client i k A linear structure is observed Reward. is the unknown model parameter to be learned by task k, and is a product with a mean of 0 and a variance of σ 2 Gaussian variable. To simplify the description and analysis, we assume that: And ||θ k ||2≤1.
[0065] Furthermore, the server-side model of the embodiment of the present invention is described as follows:
[0066] Most of the work on distributed SLB learning assumes that clients participate in collaboration without compensation, and few studies explore the rationality of client participation in collaboration, resulting in a lack of incentive mechanisms for clients in distributed SLB learning. In the server-side model, we consider the cost of client participation in collaboration and construct the following model to describe the incentive mechanism and communication process.
[0067] Distributed systems use collaboration between clients to promote learning. This process brings local learning costs, communication costs, and privacy leakage costs to clients, which hinders client participation. In our distributed learning model, we consider the collaboration cost of clients in each round of communication. In local learning with a total of T rounds, the time index set of the rounds in which communication occurs is recorded as In round In the case where the distributed system is in a communication state, client i k The unit collaboration cost for each sample in the environment will be observed The total cost is therefore Client i k Need to notify the server and To establish pricing.
[0068] In each communication round The server will select some clients Participate in collaboration and contribute to Each client i in k Assign appropriate incentives. If the server gives incentives is enough to cover the collaboration cost, then client i k It is reasonable to participate in the collaboration and upload local data. Formally speaking, when a certain incentive mechanism is used, if each participant has a non-negative profit, that is, any client participating in the collaboration satisfies Then the incentive mechanism is said to satisfy individual rationality (IR).
[0069] When designing an incentive mechanism, IR is the most basic property that needs to be met. In addition, incentive compatibility (IC) is also a property worthy of attention, which aims to resist the dishonest behavior of malicious users in the distributed system. Formally speaking, if when a certain incentive mechanism is applied, for all clients, the benefits obtained when reporting costs truthfully are the best, and reporting false costs will not make them gain more benefits, or even reduce their benefits, then the incentive mechanism is said to have the properties of IC.
[0070] Finally, the performance measurement indicators of the embodiment of the present invention are described:
[0071] In the SLB model, regret value is usually used to measure the performance of the algorithm. In the distributed learning SLB model, the goal is to minimize the cumulative regret value of all clients participating in the distributed learning task, which is defined as:
[0072]
[0073] In the above formula, Is client i k The best action that can be chosen at round t.
[0074] Reference Figure 1 and Figure 4 , Figure 1 A flowchart of a distributed SLB service recommendation method based on an incentive mechanism provided by an embodiment of the present invention, referring to Figure 1 , the method comprises the following steps:
[0075] S100, obtaining client data and building a distributed SLB model, solving it through a ridge regression estimation algorithm, and determining the amount of data uploaded by the client;
[0076] It should be noted that, in some embodiments, step S100 may include: S110, obtaining client data, and determining the historical action information of the client data and the historical reward information of the client data; S120, constructing a covariance matrix based on the historical action information of the client data and constructing a reward vector based on the historical action information and historical reward information of the client data; S130, constructing a distributed SLB model based on the covariance matrix and the reward vector, and estimating it through a ridge regression estimation algorithm to obtain an estimation result of the distributed SLB model; S140, based on an optimistic strategy, selecting the action with the highest reward confidence upper bound of the estimation result of the distributed SLB model, and obtaining the corresponding reward as the uploaded data of the client, and recording the amount of uploaded data of the client.
[0077] In some specific embodiments, in round t∈[T], active clients can perform local learning, which is equivalent to executing the classic non-distributed SLB algorithm, namely LinUCB algorithm, in the local environment under the SLB setting. Meanwhile, inactive clients only maintain their previously collected data in this round and do not perform any other operations.
[0078] The purpose of local learning is to collect historical action and reward data and use this data to learn the model θ k , thus selecting the optimal decision for this round. Specifically, t′ represents the round in which the last communication occurred. The historical data stored on the local end includes three categories:
[0079] The first type of data: round t ′The local data added later has not been uploaded for collaboration. If the client receives incentives from the server, this data will be uploaded for collaboration (aggregated into the third category of data). Otherwise, it will not be uploaded and will only be recorded in the local data to promote local learning (aggregated into the second category of data).
[0080] The second type of data: round t ′ and its previously recorded local data that has not been uploaded to participate in the collaboration;
[0081] The third type of data: round t ′ And the synchronized data previously obtained through collaboration, this part of data includes both the data uploaded by the client itself to participate in the collaboration and the data uploaded by other clients in the same learning task to participate in the collaboration, which is the source of the learning effect gain of the client.
[0082] use and S k,t Respectively represent the covariance matrices constructed from the action information in these three types of data. The specific construction forms of these matrices can be given by the following expressions:
[0083]
[0084]
[0085]
[0086] The same applies and k,t Respectively represent the reward vectors constructed from the action and reward information in these three parts of data. The construction of these vectors can be given by the following expressions:
[0087]
[0088]
[0089]
[0090] When repeating the local learning phase, client i k Use this data to learn the intrinsic model θ k The specific process is as follows. First, client i k A set of candidate actions is obtained from the environment The client hopes to predict the best action in the set to obtain the highest reward. In order to estimate the reward for each action, the client first needs to estimate the model parameters θ through ridge regression. k , client i k The estimated model parameters are denoted by Then ridge regression is used to estimate θk The specific expression is:
[0091]
[0092] in and It also includes the three types of data mentioned above. Model based on estimation The algorithm uses an optimistic strategy to select the best predicted action, and selects the action with the highest estimated reward confidence upper bound (UCB), namely:
[0093]
[0094] in is an exploration parameter that needs to be appropriately selected and is related to the confidence interval. Set to:
[0095]
[0096] Select Action Afterwards, client i k Observe the actual rewards received And update local data and and It is the key data containing the newly added information and can participate in the collaboration in the upcoming communication round. After each communication round, and is merged into the other two categories of data. and Reset to 0 respectively d×d and 0 d , to facilitate subsequent data collection and collaboration.
[0097] S200, determining the incentive cost and constructing a collaboration ratio relationship according to the amount of data uploaded by the client, and determining the client participant set;
[0098] It should be noted that, in some embodiments, step S200 may include: S210, constructing a collaborative proportional relationship based on the amount of data uploaded by the client; S220, determining the incentive cost of the client and sorting it in ascending order to obtain the sorted client incentive cost; S230, setting a threshold ratio, and dividing the clients corresponding to the collaborative proportional relationship according to the sorted client incentive cost into a set of client participants.
[0099] In some embodiments, client selection aims to optimize the collaboration effect at a limited cost. Obviously, without considering the incentive expenditure, the higher the degree of collaboration, the more learning samples it brings, which is more conducive to improving the algorithm performance. However, when considering the collaboration cost, blindly increasing the degree of collaboration may lead to excessive incentive expenditure. Therefore, we need to grasp the trade-off between the degree of collaboration and the incentive cost.
[0100] In order to measure the trade-off between the degree of cooperation and the incentive cost, we define a parameter α to represent the degree of cooperation, called the cooperation ratio, which is defined as follows:
[0101] α=P k,t / Q k,t
[0102] in represents the cumulative number of samples participating in the collaboration, Indicates the total number of samples in the current system (regardless of whether they are involved in collaboration).
[0103] In order to achieve the expected level of cooperation and thus the expected performance improvement, the cooperation ratio needs to reach a threshold ratio α h α h It is a pre-entered parameter. The server sorts the client costs in ascending order and adds the clients to the participant set in this order. Until the following conditions are met, the expression is:
[0104]
[0105] In the above formula, α represents the cooperation ratio parameter, P k,t represents the cumulative number of samples participating in the collaboration, Q k,t Represents the total number of samples in the current system, α h Represents the threshold ratio.
[0106] S300: Determine the marginal cost, and the server issues incentives to the clients in the client participant set according to the marginal cost, obtains the uploaded data of the clients, and obtains the distributed SLB service recommendation result;
[0107] It should be noted that, in some embodiments, step S300 may include:
[0108] S310, determining the marginal cost according to the custom reported cost of the client;
[0109] S320, dividing the clients into malicious clients and honest clients according to the real data cost and the custom reporting cost;
[0110] In some embodiments, the present invention adopts a metric that exposes less private information to achieve the effect of protecting privacy. Specifically, we require the client to upload the number of newly added sample data before receiving the incentive. There is no need to upload the client's historical actions and other data that are highly relevant to privacy. and To reflect the value and cost of participating in the collaboration and to seek appropriate incentives from the server. We define the reported cost as
[0111] S330, setting a communication trigger condition according to the covariance matrix of the amount of data uploaded by the client;
[0112] In some embodiments, to reduce communication overhead, communication occurs only in certain rounds, determined by a communication trigger condition, rather than in every round. The communication trigger condition is used to evaluate whether the additional local data contains enough new information to warrant a new round of collaboration, and is calculated according to the following formula, specifically:
[0113]
[0114] where αt k Indicates the time from the last communication to the current round, is the discrimination threshold. The value of will affect the performance of the algorithm. In order to ensure that the algorithm reaches a sublinear regret value, The values are as follows:
[0115]
[0116] If any client meets the communication triggering conditions, the server and client will communicate and collaborate in this round.
[0117] In some embodiments, the server calculates the appropriate incentives for the selected clients to satisfy IR and IC by determining the marginal cost for each client. As in addition to the client i k The marginal cost is defined as: Its marginal cost satisfy The marginal cost is calculated by keeping the cost reported by other clients unchanged and traversing For all cases, repeat the client selection steps so that client i k The smallest value that will not be selected by the server That is the marginal cost. Using the marginal cost, the reasonable incentive value is set as:
[0118]
[0119] In the above formula, represents a reasonable incentive value, represents marginal cost, Represents client i k The number of new samples added since the last round of communication.
[0120] Among them, the honest client reports the true cost Malicious clients report false costs
[0121] S340, based on the communication triggering condition, issuing incentives to the malicious client and the honest client according to marginal cost;
[0122] Specifically, for an honest client, if the custom reporting cost is less than the marginal cost, the honest client is selected for data service communication and receives incentives. If the custom reporting cost is greater than the marginal cost, the honest client is not selected for data service communication and cannot receive incentives.
[0123] For a malicious client, if the custom reporting cost is less than the real data cost and less than the marginal cost, or the real data cost is less than the custom reporting cost and less than the marginal cost, or the custom reporting cost is less than the marginal cost and less than or equal to the real data cost, then the malicious client is selected for data service communication and is incentivized, but the benefits brought by the incentive will not be higher than the benefits obtained when the malicious client behaves as an honest client. If the real data cost is less than the marginal cost and less than or equal to the custom reporting cost, or the marginal cost is less than the custom reporting cost and less than or equal to the real data cost, or the marginal cost is less than the real data cost and less than or equal to the custom reporting cost, then the malicious client is not selected for data service communication and cannot obtain incentives, and the benefits brought by the incentive will also not be higher than the benefits obtained when the malicious client behaves as an honest client.
[0124] In some embodiments, for i k Here is the case of an honest client:
[0125] like Since the reporting cost is less than the marginal cost, client i k Get incentives for
[0126] like Because the reporting cost is not less than the marginal cost, client i k Not selected, the incentive is 0.
[0127] For i k In the case of a malicious client:
[0128] like or Since the reported cost is still less than the marginal cost, client i k is still selected, but the incentive is still No changes;
[0129] like Since the reported cost is no longer less than the critical cost, client i k No longer selected, the incentive becomes 0, which is higher than the incentive obtained in the honest case. reduced;
[0130] like Since the reported cost is less than the critical cost, client i k Selected, motivated At this time, client i k Will obtain non-positive net income
[0131] like or Since the reported cost is still not less than the critical cost, client i k Still not selected, incentive still 0.
[0132] Therefore, the marginal cost method prevents the client from gaining more profit by deceiving the server and resists the dishonest behavior of malicious clients.
[0133] S350: If the malicious client and the honest client are selected to communicate, the covariance matrix and reward vector corresponding to the newly added local data are uploaded to the server for data communication aggregation to obtain a first service recommendation result;
[0134] S360: If the malicious client is not selected to communicate with the honest client, local data aggregation is performed on the covariance matrix and the reward vector corresponding to the newly added local data to obtain a second service recommendation result;
[0135] S370. Combine the first service recommendation result and the second service recommendation result to obtain a distributed SLB service recommendation result.
[0136] In some specific embodiments, the algorithm uses a star communication network, where N clients interact directly with a central server. In a communication round, the client selected by the server will receive sufficient incentives and participate in the collaboration. k is selected, then client i k Added local data and Upload to the server, the uploaded data will be aggregated into synchronized data.k If it is not selected, the newly added local data and Will only aggregate to local data and
[0137] The server will get the data from the client and After polymerization to S k,t and k,t , and then S k,t and k,t Synchronize to the client to participate in the learning of the local model. Because the server's goal is to facilitate the learning process of all clients, regardless of whether the client participates in the collaboration, the client will receive the updated synchronization data S sent by the server. k,t and S k,t .
[0138] In summary, the embodiments of the present invention design an incentive mechanism to motivate the participation of the client, while meeting the privacy requirements of the client before obtaining the incentive. Taking into account the model heterogeneity of distributed learning, we propose a distributed learning model based on SLB. This model introduces an incentive mechanism in distributed learning, and at the same time, it describes the differences in preferences of users of recommendation systems in different regions through a multi-task structure, and meets the privacy requirements of some institutions that do not want data to be disclosed. Based on this model, a distributed personalized recommendation algorithm is proposed to find the optimal recommendation strategy, while ensuring the client's privacy protection requirements for data such as actions and rewards until the server provides satisfactory incentives. In addition, the algorithm designs an incentive mechanism to meet IR and IC, which can prevent malicious users from defrauding higher returns by falsely reporting costs.
[0139] See also Figure 2 The embodiment of the present application also provides a distributed SLB service recommendation system based on an incentive mechanism, which can implement the above-mentioned distributed SLB service recommendation method based on an incentive mechanism. The system includes:
[0140] The first module 201 is used to obtain client data and build a distributed SLB model, and solve it through a ridge regression estimation algorithm to determine the amount of data uploaded by the client;
[0141] The second module 202 is used to determine the incentive cost and construct a collaboration ratio relationship according to the amount of data uploaded by the client, and determine the client participant set;
[0142] The third module 203 is used to determine the marginal cost. The server issues incentives to the clients in the client participant set according to the marginal cost, obtains the uploaded data of the clients, and obtains the distributed SLB service recommendation result.
[0143] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0144] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A distributed SLB service recommendation method based on an incentive mechanism, characterized in that: The method comprises the following steps: Obtain client data and build a distributed SLB model, solve it using the ridge regression estimation algorithm, and determine the amount of data uploaded by the client; Determine the incentive cost and construct a collaborative proportion relationship according to the amount of data uploaded by the client, and determine the client participant set; The marginal cost is determined, and the server issues incentives to the clients in the client participant set according to the marginal cost, obtains the uploaded data of the clients, and obtains the distributed SLB service recommendation result.
2. The method according to claim 1, characterized in that The step of obtaining client data and building a distributed SLB model, solving the model through a ridge regression estimation algorithm, and determining the amount of data uploaded by the client includes: Acquire client data, and determine historical action information of the client data and historical reward information of the client data; Constructing a covariance matrix according to the historical action information of the client data and constructing a reward vector according to the historical action information and historical reward information of the client data; Constructing a distributed SLB model according to the covariance matrix and the reward vector, and estimating it through a ridge regression estimation algorithm to obtain an estimation result of the distributed SLB model; Based on the optimistic strategy, the action with the highest reward confidence upper bound of the estimated result of the distributed SLB model is selected, and the corresponding reward is obtained as the uploaded data of the client, and the amount of uploaded data of the client is recorded.
3. The method according to claim 2, characterized in that The client data includes first-category data, second-category data and third-category data, wherein: The first type of data represents local data newly added by the client after the last data service communication. If the newly added local data has incentives, it is aggregated into the third type of data. If the newly added local data does not have incentives, it is aggregated into the second type of data. The second type of data represents local data that has not been uploaded to participate in data service communications; The third type of data represents synchronization data used for participating in data service communications.
4. The method according to claim 1, characterized in that The step of determining the incentive cost and constructing a collaborative ratio relationship according to the amount of data uploaded by the client to determine the set of client participants includes: Constructing a cooperation ratio relationship according to the amount of data uploaded by the client; Determine the incentive costs of the clients and sort them in ascending order to obtain sorted incentive costs of the clients; A threshold ratio is set, and the ranked client incentive costs are used to sequentially divide the clients that satisfy the cooperation ratio relationship into the client participant set.
5. The method according to claim 4, characterized in that The expression of the communication trigger condition is specifically as follows: In the above formula, Δt k Indicates the time from the last communication to the current round of communication, represents the discrimination threshold, represents the covariance matrix containing the uploaded data, Represents the covariance matrix of the first type of data.
6. The method according to claim 4, characterized in that The expression of the cooperation ratio relationship is specifically as follows: In the above formula, α represents the cooperation ratio parameter, P k,t represents the cumulative number of samples participating in the collaboration, Q k,t Represents the total number of samples in the current system, α h Represents the threshold ratio.
7. The method according to claim 2, characterized in that: The determining of the marginal cost, the server issuing incentives to the clients in the client participant set according to the marginal cost, obtaining the uploaded data of the clients, and obtaining the distributed SLB service recommendation result, includes: Determining marginal costs according to the custom reported costs of the client; According to the real cost of data and the custom reporting cost, the clients are divided into malicious clients and honest clients; Setting a communication trigger condition according to a covariance matrix of the amount of data uploaded by the client; Based on the communication triggering condition, issuing incentives to the malicious client and the honest client according to the marginal cost; If the malicious client and the honest client are selected to communicate, the covariance matrix and the reward vector corresponding to the newly added local data are uploaded to the server for data communication aggregation to obtain a first service recommendation result; If the malicious client and the honest client are not selected to communicate, local data aggregation is performed on the covariance matrix corresponding to the newly added local data and the reward vector to obtain a second service recommendation result; The distributed SLB service recommendation result is obtained by combining the first service recommendation result and the second service recommendation result.
8. The method according to claim 7, characterized in that The marginal cost is calculated by keeping the custom reporting cost of the remaining clients unchanged. Assuming that the custom reporting cost of the selected client changes, it is changed to the custom reporting cost of each unselected client in turn, and the client participant set is repeatedly determined so that the minimum custom reporting cost at which the selected client will not be selected by the server is the marginal cost.
9. The method according to claim 7, characterized in that: The issuing of incentives to the malicious client and the honest client according to the marginal cost includes: For the honest client, if the custom reporting cost is less than the marginal cost, the honest client is selected to perform data service communication and obtain the incentive; if the custom reporting cost is greater than the marginal cost, the honest client is not selected to perform data service communication and cannot obtain the incentive; For the malicious client, if the custom reporting cost is less than the real cost of the data and less than the marginal cost, or the real cost of the data is less than the custom reporting cost and less than the marginal cost, or the custom reporting cost is less than the marginal cost and less than or equal to the real cost of the data, then the malicious client is selected for data service communication and obtains the incentive; if the real cost of the data is less than the marginal cost and less than or equal to the custom reporting cost, or the marginal cost is less than the custom reporting cost and less than or equal to the real cost of the data, or the marginal cost is less than the real cost of the data and less than or equal to the custom reporting cost, then the malicious client is not selected for data service communication and cannot obtain the incentive.
10. A distributed SLB service recommendation system based on an incentive mechanism, characterized in that: The system comprises: The first module is used to obtain client data and build a distributed SLB model, and solve it through the ridge regression estimation algorithm to determine the amount of data uploaded by the client; The second module is used to determine the incentive cost and construct a collaboration ratio relationship according to the amount of data uploaded by the client to determine the client participant set; The third module is used to determine the marginal cost. The server issues incentives to the clients in the client participant set according to the marginal cost, obtains the uploaded data of the clients, and obtains the distributed SLB service recommendation result.
Citation Information
Patent Citations
Distributed cross-boundary service recommendation method and system based on variational reasoning
CN116361561A
Federal generalized matrix decomposition recommendation method based on incentive mechanism
CN118132845A