Worker true image discovery and recruitment method based on Shapley value

By using the Sharple value method to evaluate workers' contribution rate and trust in the group intelligence perception network, the problem of false data and system attacks caused by low credibility or malicious workers in the network is solved, and the data collection quality and platform security are improved.

CN120198093APending Publication Date: 2025-06-24CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510300593.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

There are low-reliability or malicious workers in the Clanzhi Sensing Network, resulting in false data and system attacks, affecting data quality and platform security.

Method used

Using a Sharpley value-based method, the real contribution rate is evaluated through the application quality of the worker's participation construct, and the overall trust of workers is dynamically updated, so that high-quality workers are selected for data collection.

Benefits of technology

Effectively identify and select trusted workers, improve the quality of data collection, reduce the risk of false data and system attacks, and improve the reliability and security of the platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_6
    Figure QLYQS_6
  • Figure QLYQS_17
    Figure QLYQS_17
  • Figure QLYQS_23
    Figure QLYQS_23
Patent Text Reader

Abstract

The invention discloses a worker true image discovery and recruitment method based on a Shapley value. According to the method, credible workers can be selected to collect data according to contribution values of the workers to application quality in each round. The creativity of the method is that the comprehensive credibility considers all possible worker combinations and the contribution to the total income, and the weight of the recent contribution value in calculation is larger, so that the comprehensive credibility can be dynamically updated, and a real-time effect is achieved. The method comprises the following steps that workers are selected based on candidate task information submitted by the workers, then data are collected on one hand, and the comprehensive credibility of the workers is evaluated on the other hand. And the workers with the comprehensive credibility higher than the threshold value can enter a high-credibility worker set. According to the evaluation method, data of workers with unknown credibility are compared with data of highly-credible and credible workers, and whether the collected data are true or not is identified so as to evaluate the credibility of the workers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of trusted data collection in crowd intelligence networks, and particularly relates to a method for mining the credibility of workers in crowd intelligence networks so as to recruit high-quality workers. Background Art

[0002] Crowd sensing network is a widely adopted task-centric data collection mode. A crowd sensing network mainly consists of three parts: task requesters, platforms, and workers. The process of data sensing is generally as follows: Task requesters provide their own requirements to the platform, and then the platform publishes information such as the location for data collection and the content of data collection. After receiving the task information, workers submit the set of sensing tasks they are willing to participate in, and then the platform recruits workers to participate in the data sensing of corresponding tasks and gives certain rewards to the workers. Finally, after the platform collects the data, it constructs the data into an application and provides it to the data requester. Since there are a large number of workers participating in data sensing, therefore, the crowd sensing network can obtain data for a long time, in a large range, and at low cost. Since the data submitted by workers directly affects the quality of the applications constructed by the platform, it is required that the platform selects trusted workers with high-quality data collection for data collection.

[0003] Since the perception data of workers consumes resources such as computing, storage, communication resources and time, and some tasks also need to be moved to designated locations for perception, the platform needs to pay compensation for the high cost of completing the tasks. Moreover, the higher the perception quality, the higher the requirements for workers and the more resources need to be expended. However, there are some workers with low credibility or even malicious workers in the crowd-sourced sensing network. These workers want to earn the maximum reward with the least possible cost. For this reason, these workers often do not collect data but fabricate false data reports to the platform; and some malicious workers use malicious data to launch attacks on the platform, causing the system to suffer greater losses than false data. Such things have happened during large-scale outdoor sports. For example, navigation and weather forecasts based on incorrect data can make customers lose their way and even lose their lives in harsh environments. It can be seen that there is an urgent need for effective methods to select trustworthy workers to report high-quality data to construct high-quality applications. The most decisive factor in data quality is the authenticity of the data, which requires that the data reported by workers be consistent with the real data, also known as truth discovery or real data discovery. There have been some studies on real data discovery, which mainly obtain real data from the acquired data through computational methods, mainly a class of methods based on mathematical calculations such as the average method, the median method, the weighted average method, etc. The main feature of this kind of method is to select n workers to perceive the same perception target at the same time. Then, the average value, median value, and weighted average of these n are obtained as the real data. The basic idea behind using this kind of method is that most workers in the crowd-sourced sensing network are trustworthy and the data reported by trustworthy workers are all real. It can be seen that in a real crowd-sourced sensing network, it is difficult to even obtain the truth or falsehood of the data reported by workers. In view of the above research status, in this paper, we use the Shapley value calculation method to further calculate the contribution rate of workers to the task according to the completion degree of the applications constructed by the workers. Since the final quality of the application is the result of the influence of all factors, using the contribution rate of workers to the quality of the finally formed application to evaluate the contribution rate of workers to the task is more comprehensive than using one or several factors to represent the contribution rate in previous studies. The method of the present invention is to propose an effective method to identify the contribution rate of workers, so as to select trustworthy workers to perceive data, and thus high-quality crowd-sourced sensing applications can be constructed. Summary of the Invention

[0004] The present invention discloses a method for discovering and recruiting real workers based on the Shapley value. Aiming at the situation of low credibility in data collection in the crowd intelligence network, or malicious workers submitting false or malicious data, the inventive method proposes a method for selecting credible workers to collect data based on the comprehensive trustworthiness of workers. Workers are individuals who hold sensing devices such as mobile phones independently of the platform, and obtain rewards by sensing data and submitting data. Although this method can obtain a large amount of data quickly and at low cost, since it is difficult for the platform to verify whether the data submitted by workers is real data, it has troubled the development of crowd intelligence network applications. Therefore, the present invention proposes a method for judging the real contribution of workers through the applications constructed by workers, and then dynamically updating the comprehensive trustworthiness of workers to select high-quality workers to sense data to achieve the above goal.

[0005] In the cooperation of a task, the completion rate of the total task is often not simply the sum of the contribution rates of multiple workers. Therefore, we cannot simply regard the increment of the contribution rate of a single worker to the total task as its real contribution rate, but need to find an algorithm to quantify the real contribution rate of each worker during cooperation. This algorithm is usually used for the benefit distribution of each member in cooperation, and here we use it to determine the contribution rate of each worker in the task. The Shapley value is a distribution method derived from game theory, which is specifically used to calculate the marginal contribution of each worker to the overall result in cooperation.

[0006] The method proposed by the present invention is based on the fact that the contribution rate of an initial single worker can be obtained by recruiting workers and interacting with them. Then, when multiple workers are selected to jointly participate in the same data collection task, on the one hand, the quality of the collected data can be improved, and on the other hand, the real contribution of the workers can be evaluated. The method is to calculate the marginal contribution when the same worker joins different worker combinations and then take the average value, so as to effectively identify the real contribution rate of the worker in this task and evaluate its trustworthiness. When the number of credible workers obtained by the system is small, the system automatically increases the number of unknown recruits, thereby increasing the number of times of sensing the contribution rate of workers and obtaining a more real contribution rate of workers. When the number of credible workers obtained by the network reaches an available state, known high-quality workers are actively recruited to save costs. Finally, the purpose of improving the quality of data collection and saving costs is achieved.

[0007] The technical solution of the invention is as follows:

[0008] 1. A method for discovering and recruiting real workers based on the Shapley value, characterized by including the following steps:

[0009] (1) The platform recruits staff to sense data and processes the sensed data to construct an application Application Can be divided into multiple subroutines a and constructed in multiple rounds; Application Construction round For the application subroutine a, it can be represented by M a,t to represent the task set that constitutes the application program a in the t-th round, where The worker set is defined as

[0010] (2) Initialization: S is the total combination of recruited workers, and the worker combination s is a subset of S. For each worker combination s, initialize the task completion rate v(s)=0; for each worker w i , initialize the cooperation contribution For application a, the high-trust worker set and the low-trust worker set are initialized as Set the threshold for entering the high-trust worker set as The threshold for entering the low-trust worker set is Initialize the probability ε of selecting a new worker combination to be 0.5;

[0011] (3) When the round number is reached, for each application a, perform the following steps:

[0012] {

[0013] 1) The system platform publishes the task set M that requires collecting data from |M a,t | locations. After the workers in the sensing network obtain the task set M a,t for collecting data, there are a,t workers applying for data collection and submitting candidate task information to the platform, where where represents the k-th task that worker w i is willing to participate in during the t-th round;

[0014] 2) After the system platform receives the applications, perform the following operations on each candidate task information :

[0015] If worker w i is in the high-trust worker set , recruit workers in descending order of the comprehensive trust degree in each round, and successively select the combination S = S ∪ {w i} to participate in the task

[0016] If the number of the worker set S has not reached the threshold |M a,t | at this time, then select the remaining worker w i to participate in this sensing task Data collection, S = S ∪ {w i};

[0017] If the number of the worker set S reaches the threshold |M a,t |, then a new untested worker is selected with probability ε, and a worker with the highest task completion rate v(w i ) that has been evaluated is selected with probability 1 - ε;

[0018] For each additional worker w i , the platform or the task applicant evaluates the task completion rate v(S). For example, it is set to 0.95 if the application quality is high and 0.70 if the application quality is low. In practice, it can be set according to needs; then update the parameter ε. If it is greater than 0.1, reduce this value, such as reducing ε by 0.02 per round;

[0019] 3) For the case where the task completion rate is not obtained in some combinations s, estimate the task completion rate according to the second - order interaction formula:

[0020]

[0021] where v(w i ) represents the independent contribution of a single worker, and β ij represents the interaction effect between worker w i and w j . The interaction effect β ij is calculated according to the following formula:

[0022] β ij = v(w i , w j ) - v(w i ) - v(w j )

[0023] 4) Then calculate the weighting factor ω(|s|) for each subset s of worker combinations. The role of this weight is to ensure that all possible joining orders are considered, and the calculation method is as follows:

[0024]

[0025] where |s| represents the number of workers included in the set s, and (|s| - 1)! represents that there are (|s| - 1)! sorting orders when worker w i participates in the set, and the remaining (|M a,t | - |s|) members have (|M a,t | - |s|)! sorting orders;

[0026] 5) Then, according to the recorded task completion rate v(s) of the worker combination s and the task completion rate v(s\{w i after removing worker w i}), combined with the Shapley method to evaluate the true contribution of its workers The selection method is as follows:

[0027]

[0028] represents the true contribution of worker w i in the set; where s represents a subset of set S, and v(s\{w i ) represents the task completion rate that the set can obtain after removing w i , ω(|s|) is a weighting factor, and the role of this weight is to ensure that all possible joining orders are considered. The calculation formula of the weighting factor is as follows:

[0029] 6) Using the above, we can obtain Normalize the Shapley values of different workers to obtain the contribution rate of workers in the task. The normalization formula is as follows:

[0030]

[0031] where C i,j (t) represents the contribution rate of worker w i to task t j , represents the k-th task option submitted to the platform in the t-th round is the task in the composition of ;

[0032] 7) When obtaining the contribution rate of workers, we calculate the weight decay function h(t) to assign different weights to the contribution rates C i,j (t) of the current and previous rounds. The weight decay function h(t) is as follows:

[0033]

[0034] where is the decay coefficient, which controls the weights of different rounds. The earlier the weight, the lower it is, and it can take 0.9; when is relatively low, the weight h(t) will decrease rapidly; this means that the algorithm assigns rapidly decreasing weights to the data in the early rounds, so as to focus on the most recent rounds faster. Then the comprehensive trustworthiness i of worker w is calculated as follows:

[0035]

[0036] where is the weight factor of the current round t. Through this weighted average formula, we can obtain the comprehensive trustworthiness of each worker in the current round t;

[0037] 8) After obtaining the comprehensive trustworthiness if the comprehensive trustworthiness of the worker is greater than the high threshold it enters the high-trust worker set In the high-trust worker set, sort the workers according to their comprehensive trustworthiness from high to low; if the comprehensive trustworthiness of the worker is lower than the threshold it enters the low-trust worker set The classification formulas for high-trust and low-trust workers are as follows:

[0038]

[0039] }

[0040] Beneficial effects

[0041] The present invention discloses a method for discovering and recruiting true workers based on the Shapley value. The basic idea of the inventive method is as follows: in an actual crowdsensing network, a task is often completed by a combination of multiple workers, and the completion rate of a constructed application is not simply the sum of the contribution rates of multiple workers. Moreover, due to the existence of malicious workers, a worker selection scheme based on the true contribution of workers when completing tasks is required. The crowdsensing network platform initially obtains the individual contributions of some workers through actual interactions. Then, when selecting workers, a part of trusted workers and unknown workers are actively selected to jointly complete tasks, and their contribution rates are continuously obtained. The purpose of selecting highly trusted workers is to ensure that the platform collects real data and to achieve good results for the current system platform. The reason for selecting a part of workers with unknown trust levels is that: in many cases, it is forced to select workers with unknown trust levels, which may result in the selection of low-trust or malicious workers, thereby damaging the system. Therefore, some workers with unknown trust levels need to be selected to collect data simultaneously and participate in constructing applications together with trusted workers to evaluate the true contributions of unknown workers participating in new tasks. Thus, through the selected workers with uncertain trust levels, if the comprehensive credibility of the selected workers is higher, it can provide a long-term available worker basis for subsequent tasks and improve the selection quality. If the comprehensive credibility of the selected workers is lower, they will be added to the set of low-trust workers to avoid repeated selection, so it will only affect the current result, thereby ensuring the robustness of the method. After this process continues, more and more workers that the platform can identify will emerge. When the number of trusted workers that the platform can identify reaches a certain amount, the platform can select workers from the set of highly trusted workers to improve the quality of the constructed application. At this time, the platform achieves the balance between the exploration of unknown workers and the utilization of known workers by controlling ε: when there are not enough trusted workers, the number of Shapley tests is increased to ensure quality; when the number of trusted workers is sufficient, the number of exploratory tests is reduced to save costs. Brief Description of the Drawings

[0042] Figure 1 It is for the process of verifying the credibility of workers participating in sensing tasks;

[0043] Figure 2 It is for the comparison between the method of the present invention and the greedy selection method under different ε configurations; Detailed Embodiment

[0044] To facilitate the understanding of the present invention, the present invention will be described more comprehensively and meticulously below in conjunction with the accompanying drawings of the specification and preferred embodiments. However, the protection scope of the present invention is not limited to the following specific embodiments.

[0045] Unless otherwise defined, all professional terms used hereinafter have the same meaning as commonly understood by those skilled in the art. The professional terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the protection scope of the present invention.

[0046] Unless otherwise specified, various raw materials, reagents, instruments, and equipment used in the present invention can be obtained through market purchase or prepared by existing methods.

[0047] Examples:

[0048] In a smart city, to sense data on the environment in the city, such as traffic flow, the platform can issue sensing tasks. At this time, a large number of workers in the city can receive the tasks, perform data sensing, and send the data to the platform. However, there are some untrusted users among the workers who upload virtual or malicious data, making the applications on the platform unreliable. At this time, the following method can be used to obtain and evaluate the trustworthiness of referents in real time and efficiently, so that the platform can obtain high-quality real data.

[0049] The experimental results of the inventive method are given below.

[0050] Figure 1 The experimental results were obtained in the following experimental environment. The system consists of a platform and workers. The platform provides services to task requesters by constructing high-quality application programs, and thus obtains. An application program contains multiple subroutines, which need to be implemented by different combinations of workers. We randomly selected workers as sensors in each dataset, and the number of sensing workers ranged from 50 to 200. For simplicity, we assume that |M a,t | is consistent, that is, |M a,t | = |M|. The number of subroutines |M| in each application program ranged from 5 to 20, that is, the number of workers recruited. We analyzed the inventive method for the identification of different types of worker groups, especially the identification of highly trustworthy worker groups . The experimental results Figure 1 are shown as follows, Figure 1 which shows the process of identifying highly trustworthy workers in the trust exploration of the inventive method. The inventive method explores the trustworthiness of different worker groups and classifies the worker groups by dynamically adjusting the worker exploration and worker selection strategies, ensuring the rationality of worker recruitment and the efficient completion of tasks. In the test, the worker trust detection and discovery efficiency that the inventive method can ultimately achieve is as high as 94.79%. For the simulated dataset, the number of finally developed low-trust workers accounts for 18.51% of the total number , and the number of highly trustworthy workers that can be maintained accounts for 42.86% of the total number .

[0051] Figure 2It is a comparison between the method of the present invention under different ε configurations and the method of the comprehensive exploration index without ε control. For the convenience of comparison, we set different ε = 0.1, 0.2, 0.3, and compared the influence of the method of the present invention under different ε and the greedy method without the comprehensive exploration index on worker recruitment. From Figure 2 It can be seen that the results are as follows: Under different ε parameters, the method of the present invention is superior to the greedy method, and the high ε configuration helps to quickly screen out highly credible workers. However, it can be seen from the figure that the cost cannot be well controlled under high ε. The reason is that the overly greedy selection confuses some highly credible workers and low-credible workers, and at the same time affects the income and exploration cost of the current round. Specifically: The relative frequencies of workers with quotes below 0.4 being recruited under the configurations of ε = 0.3, ε = 0.2, and ε = 0.1 are 2.72 times, 2.75 times, and 2.85 times that of the greedy method respectively. The relative frequencies of workers with Shapley values above 0.4 being recruited under the configurations of ε = 0.3, ε = 0.2, and ε = 0.1 are 2.52 times, 2.19 times, and 2.09 times that of the greedy method respectively. Generally speaking, the configuration of ε = 0.1 is superior in controlling the task cost, while ε = 0.3 has more advantages in recruiting high-quality workers.

Claims

1. A worker truth discovery and recruitment method based on Shapley value, characterized in that The following steps are involved: (1) The platform recruits workers to perceive data and process the perceived data to build applications app Can be divided into multiple subprograms a, multiple rounds of construction; application The build round For application subroutine a, you can use M a,t represents the set of tasks that constitute application a in round t, where The worker set is defined as (2) Initialization: S is the total number of recruited worker combinations, worker combination s is a subset of S, for each worker combination s, the initialization task completion rate v(s) = 0; for each worker w i , Initialize collaborative contributions For application a, a set of highly trusted workers and the set of low-trust workers Initialize to The threshold for entering the high-confidence worker set is set to The threshold for entering the low-trust worker set is Initialize the probability of selecting a new worker combination ε=0.5; (3) Number of rounds When , for each application a, perform the following steps: { 1) The system platform releases the data that needs to be collected a,t |Task set M for location data a,t , the workers in the sensing network get the task set M of collecting data a,t After that, there Workers apply for data collection and submit candidate task information to the platform in Represents worker w i The kth task that the participants are willing to participate in in the tth round; 2) After receiving the application, the system platform will Do the following: If the worker w i In the high-trust worker collection In each round, workers are recruited from high to low comprehensive trust, and combinations S = S ∪ {w i }Participate in the task If the number of workers in the set S does not reach the threshold value |M a,t |, then select the remaining workers w i To participate in this perception task Data collection, S = S ∪ {w i }; If the number of workers in the set S reaches the threshold |M a,t |, then select a new untested worker with probability ε, and select the evaluated worker with the highest task completion rate v(w i ) Each additional worker w i , the platform or task applicant evaluates the task completion rate v(S). For example, v(S) of high-quality applications is set to 0.95, and v(S) of low-quality applications is set to 0.

70. In practice, it can be set as needed. Then update the parameter ε. If it is greater than 0.1, reduce the value, such as reducing ε by 0.02 per round. 3) For the case where the task completion rate is not obtained in some combinations s, the task completion rate is estimated according to the second-order interaction formula: Where v(w i ) represents the independent contribution of a single worker, β ij Represents worker w i and w j The interaction effect, β ij Calculated according to the following formula: β ij =v(in i ,In j )-v(in i )-v(in j ) 4) Then, for each worker combination subset s, a weighting factor ω(|s|) is calculated. The role of this weight is to ensure that all possible joining orders are considered. The calculation method is as follows: Here, |S| represents the number of workers contained in set s, and (|s|-1)! represents worker w i There are (|s|-1)! sortings when participating in the set, and the remaining (|M a,t |-|s|) members have (|M a,t |-|s|)! sort; 5) Then based on the recorded task completion rate v(s) of worker combination s and the removal of worker w i The task completion rate v(s\{w i }), combined with the Shapley method to evaluate the real contribution of its workers The selection method is as follows: Represents worker w i The real contribution in the set; where s represents a subset of the set S, v(s\{w i }) means removing w i The task completion rate that the set can obtain after that, ω(|s|) is the weighting factor. The role of this weight is to ensure that all possible joining orders are considered. The calculation formula of the weighting factor is as follows: 6) Using the previous one, we can get The Shapley values ​​of different workers Normalize to get the worker's contribution rate in the task. The normalization formula is as follows: Among them C i,j (t) represents the worker w in round t i For task t j The contribution rate, represents the kth task option submitted to the platform in round t is in the composition tasks in 7) When we get the worker's contribution rate, we calculate the weight decay function h(t) to calculate the contribution rate C of the current and previous rounds i,j (t) is assigned different weights, and the weight decay function h(t) is as follows: in is the attenuation coefficient, which controls the weights of different rounds. The earlier the weight, the lower it is. It can be taken as 0.

9. When t is low, the weight h(t) decreases rapidly; this means that the algorithm quickly reduces the weight given to data from earlier rounds, focusing more quickly on the most recent rounds, and then the worker w i The overall trust The calculation formula is as follows: in is the weight factor of the current round t. Through this weighted average formula, we can get the comprehensive trust of each worker in the current round t; 8) After obtaining comprehensive trust After that, workers' comprehensive trust Greater than high threshold Entry In the high-trust worker collection According to the comprehensive trust of workers Ranking from high to low; comprehensive trust of workers Below threshold Enter the low-trust worker set The classification formula for high-trust and low-trust workers is as follows: }。