Credible method for improving data quality and completion rate based on path dependence

Through the combination of trust mechanism and path dependence principle, trusted workers are identified and reward strategies are dynamically adjusted, which solves the problems of low data quality, low task completion rate and high cost in group intelligence perception, and improves data quality and task completion rate and reduces costs.

CN120338427APending Publication Date: 2025-07-18CENT SOUTH UNIV +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510522207.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

There are problems in group intelligence perception technology with low data quality, low task completion rate and high cost. The existing incentive mechanism is difficult to meet the needs of high data quality, high task completion rate and low cost at the same time.

Method used

Introduce trust mechanism and path dependence principle, classify workers as trustworthy workers and workers with uncertain trust, use trustworthy evolution mechanism to identify trustworthy workers and prioritize their data, combine the path dependence principle of behavioral economics, and dynamically adjust reward strategies to motivate workers to complete non-hot-spot tasks, forming path dependence to reduce costs.

Benefits of technology

The data quality has been improved by 50.54%, the task completion rate has reached 41.94%, and the cost has been reduced by 44.09%, achieving balanced optimization of data quality, task completion rate and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338427A_ABST
    Figure CN120338427A_ABST
Patent Text Reader

Abstract

The invention discloses a credible method for improving data quality and completion rate based on path dependence, and relates to the field of data collection of a crowd-sourcing network. Constructing a worker set, wherein an initial worker set comprises initial credible workers and workers with undetermined credibility; the platform issues tasks to be executed, wherein the tasks to be executed are divided into hotspot tasks and non-hotspot tasks; making a reward strategy according to the task completion condition of the worker and the execution task type; workers select intention tasks according to own conditions and submit the intention tasks to the task distribution platform; the platform selects a worker and notifies the selected worker to do a task, the worker submits data to the platform after sensing the data, and the platform obtains the data; and the platform updates the credibility of the workers according to the task completion conditions of the workers, and pays a reward. According to the method, the problems that the data quality is difficult to guarantee, the task completion rate is insufficient and the cost is too high are effectively solved, so that the balance optimization of data quality, completion rate and cost control is realized while the requirements of different application scenes are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data collection in a crowd intelligence network, and more specifically to a trusted method for improving data quality and completion rate based on path dependence. Background Art

[0002] Mobile Crowdsensing (MCS), as an emerging data collection technology, utilizes mobile devices equipped with sensors and a large number of user groups to complete sensing tasks, and has shown great application potential in multiple fields such as military, industry, life health, and environmental monitoring. However, the widespread application of the crowdsensing technology faces three major problems: low data quality, low task completion rate, and high cost.

[0003] Data quality is one of the core issues of MCS. High-quality data needs to be consistent with real data. However, due to the inevitable existence of malicious workers, they may submit false data for their own interests, which makes it difficult to guarantee data quality. Low data quality will make it difficult for the crowdsensing network to develop.

[0004] The task completion rate refers to the proportion of tasks published by the platform that are completed. A high completion rate is very important for many crowdsensing applications. If some sensing tasks cannot be completed and the completion rate is low, there will be blanks in the area that needs to be monitored by the entire application, resulting in low quality of such applications and affecting their development. However, workers tend to execute tasks based on opportunism, preferring tasks with low cost and easy-to-complete tasks, resulting in no one being interested in remote or difficult tasks. Although increasing task rewards can motivate workers to execute non-popular tasks, this method will significantly increase the platform cost.

[0005] How to control costs is also another key issue in the crowd intelligence network. The platform needs to pay workers' remuneration to compensate for the cost of their sensing tasks. However, the platform hopes to obtain high-quality data and a high task completion rate at the lowest cost, while workers expect to receive higher remuneration. There is an interest conflict between the two. Existing incentive mechanisms such as the auction mechanism can achieve market equilibrium, but there are still deficiencies in terms of data quality and completion rate. In addition, incentive mechanisms based on behavioral economics can reduce costs, but they ignore the possible dishonest behavior of workers and the uncertainty of data quality.

[0006] The existing technologies have obvious deficiencies in simultaneously meeting high data quality, high task completion rate, and low cost. Therefore, there is an urgent need for an innovative mechanism that can comprehensively solve the above problems to promote the widespread application of MCS technology. Summary of the Invention

[0007] In view of this, the present invention provides a reliable method for improving data quality and completion rate based on path dependence. By innovatively introducing a trust mechanism and the theory of modern economics, this method can motivate workers to provide high-quality data while increasing the completion rate of tasks and significantly reducing the operating costs of the platform. Compared with the prior art, the present invention effectively solves the problems of difficult guarantee of data quality, insufficient task completion rate, and excessive costs in previous methods, thereby achieving a balanced optimization of data quality, completion rate, and cost control while meeting the requirements of different application scenarios. Generally speaking, the present invention first classifies workers by the platform, including trusted workers and workers with undetermined trust levels, and assigns a trust level of 1.0 to trusted workers and a trust level of 0.5 to workers with undetermined trust levels. The main points of the method of the present invention are as follows: First, a mechanism of trusted evolution is used to identify the trust levels of workers. When selecting workers, trusted workers are selected according to their trust levels, and the data submitted by trusted workers is used as the basis for calculating the true value of the data, thereby improving data quality. Second, the principle of path dependence in behavioral economics is used to reduce costs and increase the task completion rate. The main idea is as follows: The reward mechanism given by the platform to workers is that if a worker completes a hot task, a certain amount of remuneration is given to the worker. However, if a worker completes a non-hot task, in addition to the normal reward, an additional reward with a certain probability is also given, so as to motivate workers to do non-hot tasks, thereby increasing the task completion rate. However, if only this is done, the cost of the platform will be high. Therefore, the present invention uses the mechanism of path dependence to increase the task completion rate while minimizing the remuneration paid by the platform. The method is as follows: In the initial stage, the probability of additional rewards given by the platform to workers who complete non-hot tasks is higher than the announced probability of additional rewards, so that workers will think that the remuneration for completing non-hot tasks is high, and thus actively complete non-hot tasks. After such a period of time, workers will think that the remuneration for completing non-hot tasks is high, and thus always choose non-hot tasks to do. This forms a path dependence for workers. After workers form a path dependence, the platform reduces the probability of additional rewards, thereby reducing expenditures and recovering the costs paid in the early stage, so as to reduce the overall cost. Although workers will later find that the remuneration for completing non-hot tasks has decreased, the path dependence will still persist for a period of time, and this period of time can recover the costs. The entire process achieves the goal of increasing the task completion rate and reducing costs. Coupled with the identification and selection of trusted workers, the simultaneous optimization of quality, cost, and task completion rate is achieved.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] A reliable method for improving data quality and completion rate based on path dependence, characterized by comprising the following steps:

[0010] Construct a set of workers. The initial set of workers includes initial trusted workers and workers with undetermined trust levels;

[0011] The platform publishes tasks to be executed, and the tasks to be executed are divided into hot tasks and non-hot tasks;

[0012] Formulate a reward strategy based on the task completion situation of workers and the type of tasks executed;

[0013] Workers select intended tasks according to their own situations and submit them to the task assignment platform;

[0014] The platform selects workers and notifies the selected workers to perform tasks. After the workers perceive the data, they submit the data to the platform, and the platform obtains the data;

[0015] The platform updates the trust level of workers according to the task completion situation of workers and pays remuneration.

[0016] Optionally, the initial set of workers is U, the initial set of trusted workers is U A , with a trust level of 1, and the workers with undetermined trust levels are U B , with an initial trust level of 0.5, and the trust level is between [0, 1], where 0 means untrusted.

[0017] Optionally, the platform publishes tasks to be executed Task t i contains the following information: where {x i , y i} represents the coordinates of the task, represents the budget; the platform divides the tasks into a hot task set based on the historical completion situation of the tasks: Non-hot task set: Tasks to be executed Hot tasks refer to tasks with an execution intention higher than a preset value and a completion probability higher than a preset value; non-hot tasks refer to tasks with an execution intention lower than a preset value and a completion probability lower than a preset value.

[0018] Optionally, formulate a reward strategy according to the task completion situation of workers and the type of tasks executed. Specifically: The platform formulates the payment to workers as shown in Equation (1), represents the remuneration paid to worker i for completing task k in the j-th round of data collection announced by the platform; to encourage workers to participate in more tasks in set T np Therefore, if the worker participates in tasks in T p in this round, the remuneration received by the worker is only the fixed price B; if the worker participates in tasks in T np in this round, then on the basis of the fixed price B, there is a probability of α disp to obtain an additional reward, αdisp is the winning probability publicly shown by the platform to the workers; d i,k represents the distance between worker i and task k. When d i,k = 0, the moving cost of the worker should be 0; δ p represents the moving cost coefficient per unit distance. Additionally, the true probability of a worker obtaining an extra reward set by the platform: α DT , which is not publicly announced by the platform;

[0019]

[0020] The tasks selected by the worker specifically include: the worker traverses all the tasks within his own moving range, calculates the corresponding R i,j,k (·) for each task, and then selects max(R i,j,k (·)) as the task willing to do in the j-th round; The definition of R i,j,k (·) is as follows:

[0021]

[0022] Among them, represents the expected utility of the worker for this task in this round; is calculated as in formula (3); d i,k represents the distance between worker i and task k; D i,j represents the degree of path dependence of the worker in this round. The deeper the worker's degree of path dependence on executing , the more inclined the worker is to select tasks; D i,j is calculated as in formula (14); ω U , ω d , ω D are parameters representing the parameters for expected utility, distance, and degree of path dependence;

[0023]

[0024]

[0025] The cost of the worker is shown in equation (4). The cost for the worker to participate in the task consists of 2 parts, namely the fixed data collection cost A and the moving cost d i,k represents the distance between the worker and the task; When d i,k = 0, the moving cost of the worker should be 0. Therefore, the moving cost is set as δ c represents the moving cost coefficient per unit distance, represents the reward estimated by the worker himself, Denote the reward probability estimated by the worker himself. The probability of obtaining an additional reward is related to α announced by the platform disp .

[0026] Optionally, the method for the platform to select workers is as follows:

[0027] For any task, if there is at least 1 worker among the candidates who belongs to the initial trusted worker set, only 1 worker from the trusted worker set is recruited, and the remaining κ0 - 1 workers are recruited from the set U B ; if none of the workers who declared this task in this round belong to U A workers, then only 1 worker is recruited for this task; where κ0 is the number of recruited workers, and the value range is 3 - 20.

[0028] Optionally, the platform updates the worker trust degree according to the task completion situation of the worker. Specifically: the factor that determines whether the trust degree rises or falls is the matching rate between the worker data and the benchmark data, where the benchmark data refers to the data received by the platform submitted by worker j who belongs to the set U A , and d i,j,k is the data submitted by the worker who belongs to the set U B ; denote a i,j,k as the accuracy rate, and τ a as the accuracy threshold for accepting this data; when the accuracy rate reaches the threshold set by the platform, it is considered that the current worker's data collection behavior for this data is normal, and the trust degree is increased for him, otherwise it is decreased; the formulas for increasing and decreasing the trust degree are as shown in Equation (8):

[0029]

[0030] denotes the trust degree of worker i for task k in the j - th round, θ T denotes the parameter, denotes the maximum trust degree.

[0031] Optionally, the specific steps for paying the remuneration are as follows:

[0032] The platform defines two stages for workers, one is the investment stage, denoted as the DT stage, and the other is the harvesting stage, denoted as the DM stage; workers initially belong to the DT stage. When the worker meets the following requirements, it is determined that he enters the DM stage and recovers the cost in this stage, otherwise the worker is in the DT stage; the worker's deviation prediction for α disp is where γ represents the deviation coefficient; when , it is considered that the worker has formed a prediction deviation for the small - probability event, is calculated according to Equation (12) or Equation (16) respectively according to the stage where the worker is located; when At this time, it is considered that the worker has reached the path dependence degree required by the platform for participating in the non-hot task T np ; The path dependence degree evaluation function is as follows:

[0033]

[0034] Wherein, represents the set of all completed T np tasks by worker i in round j; τ D is the worker path dependence degree threshold, which is set by the platform; The calculation of is as shown in formula (3);

[0035] If the worker belongs to the DT stage: In the DT stage, the idea of the invention method is to set a higher α r by the platform, np increase the actual reward of the worker for completing T np in the DT stage, and then increase the expected utility of completing T

[0036] The platform sets the actual winning probability α DT ; α DT is higher than the publicly announced probability α disp , so that the posterior probability of the worker's actual winning is always higher than his estimated probability Under this setting, the worker will continuously increase the expected probability of obtaining the additional reward for T np , and the expected utility of the worker for completing T np will also increase continuously accordingly; The increase in expected utility encourages the worker to participate more actively in T np tasks, strengthening the path dependence on participating in T np ; At the same time, it also improves the task completion rate of the system and reduces the proportion of T np ;

[0037] Therefore, the worker's estimated probability is a function that is positively correlated with the number of executions of T np ; And due to the marginal diminishing effect, as continues to increase, its rising speed continues to decline and finally approaches the true probability α infinitely; Therefore, it is defined as DT as follows: ;

[0038]

[0039] The actual reward given by the platform to the worker is

[0040]

[0041] If the worker belongs to the DM stage and enters the DM stage, it means that the cultivation of path dependence and probability estimation deviation has been completed. When the worker has formed a path dependence on the execution of T np The platform maintains α disp unchanged at this stage and reduces the actual winning probability of the worker from α DT to α DM , where α DT >α disp >α DM , thus reducing costs;

[0042] The calculation of the expected probability of the worker in the DM stage is set as follows:

[0043]

[0044] St:α DM <α DT +[π(α disp ) - α disp

[0045] At the beginning, due to path dependence and thinking inertia, the worker will be more inclined to use the expected probability cultivated in the DT stage as the expected probability in the DM stage. Therefore, it is closer to As the actual probability of the platform in the DM stage is lowered, the posterior probability of the worker winning the award continually decreases, and the impact on the worker gradually increases, and finally approaches the posterior probability And because of the platform control Therefore it will approach α DM infinitely;

[0046]

[0047] The actual reward given by the platform to the worker is

[0048]

[0049] ​As can be seen from the above technical solutions, compared with the prior art, the present invention provides a reliable method for improving data quality and completion rate based on path dependence, aiming to comprehensively solve the three key problems of data quality, task completion rate, and cost reduction in the crowd-sensing network. By introducing a trust score mechanism and a dynamic reward adjustment strategy, the present invention can screen out reliable workers and encourage them to submit high-quality data, while improving the task completion rate and controlling costs. The experimental results show that the present invention performs excellently in terms of data quality, and the data quality of the method of the present invention has increased by 50.54%. In terms of the task completion rate, the task completion rate of the method of the present invention is as high as 41.94%, and the cost has been reduced by 44.09%. In summary, the present invention has achieved optimization in multiple key dimensions, providing an efficient, reliable, and innovative solution for the wide application of MCS technology, and having important practical application value and popularization prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0051] Figure 1 It is the change of trust for different types of workers;

[0052] Figure 2 It is the distribution of task completion;

[0053] Figure 3 Task completion rate;

[0054] Figure 4 It is the flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0056] The embodiment of the present invention discloses a reliable method for improving data quality and completion rate based on path dependence, as Figure 4 shown, which includes the following steps:

[0057] Construct a set of workers. The initial set of workers includes initial reliable workers and workers with undetermined trust scores;

[0058] The platform releases tasks to be executed, and the tasks to be executed are divided into hot tasks and non-hot tasks;

[0059] Formulate a reward strategy based on the task completion situation of workers and the types of tasks executed;

[0060] Workers select intended tasks according to their own situations and submit them to the task allocation platform;

[0061] The platform selects workers and notifies the selected workers to perform tasks. After the workers perceive the data, they submit the data to the platform, and the platform obtains the data;

[0062] The platform updates the trust level of workers according to the task completion situation of workers and pays remuneration.

[0063] Specifically, it includes the following steps:

[0064] (1) The set of workers is U. Initially, a small proportion of workers are trusted workers, and the initial set of trusted workers is U A , whose trust level is 1, and the remaining is the set of workers with undetermined trust levels U B , whose initial trust level is 0.5, and the trust level is between [0, 1], where 0 means untrusted and 1 means high credibility;

[0065] (2) The platform releases tasks that need to be executed For each task, such as task t i mainly contains the following information: where {x i , y i} represents the coordinates of the task, represents its budget. The platform divides the tasks into a hot task set according to the historical completion situation of the tasks: or non-hot tasks: Hot tasks refer to tasks that are willing to be executed by a large number of workers, so they are relatively easy to complete and have a high probability of being completed; non-hot tasks refer to tasks that are rarely willing to be executed by workers due to reasons such as remoteness, and have a low probability of being completed. Therefore, to improve the task completion rate, workers need to be motivated to do non-hot tasks;

[0066] (3) The platform formulates a reward strategy: The platform formulates the remuneration paid to workers As shown in formula (1), represents the remuneration given to worker i for completing task k in the j-th round of data collection announced by the platform; to encourage workers to participate in more tasks in set T np Therefore, if the task participated by the worker in this round is in T p in the tasks, the remuneration received by the worker is only the fixed price B; and if the task participated by the worker in this round is in T np task, then on the basis of the fixed price B, there is αdisp The probability of obtaining an additional reward is α disp which is the winning probability publicly shown by the platform to the workers; d i,k represents the distance between worker i and task k. When d i,k = 0, the moving cost of the worker should be 0; δ p represents the unit distance moving cost coefficient. Additionally, the real probability of the worker obtaining an additional reward set by the platform: α DT which the platform does not publicly announce.

[0067]

[0068] (4) Worker selects a task: The worker traverses all tasks within their moving range and calculates the corresponding R i,j,k (·) for each task in this round, and then selects max(R i,j,k (·)) as the task the worker intends to do in the j-th round. The definition of R i,j,k (·) is as follows:

[0069]

[0070] where, represents the expected utility of the worker for this task in this round. The higher the expected utility, the more inclined the worker is to select the task. is calculated as in formula (3); d i,k represents the distance between worker i and task k. The farther the distance, the lower the worker's evaluation of the task, and the less inclined the worker is to select the task. D i,j represents the degree of path dependence of the worker in this round. The deeper the worker's path dependence on executing , the more inclined the worker is to select task. D i,j is calculated as in formula (14). ω U , ω d , ω D are parameters representing the parameters for expected utility, distance, and degree of path dependence;

[0071]

[0072] The cost of the worker is as shown in formula (4). The cost for the worker to participate in a task consists of 2 parts, namely the fixed data collection cost A and the moving cost d i,k represents the distance between the worker and the task. The farther the distance, the exponentially increasing the moving cost of the worker. When d i,k = 0, the moving cost of the worker should be 0, so the moving cost is set to δ c represents the unit distance moving cost coefficient, denotes the reward estimated by the worker himself denotes the probability of reward estimated by the worker himself, which is the result estimated from the historical data of his actual winning of the award. The probability of his obtaining an additional bonus amount is not necessarily equal to α announced by the platform disp , α DT is not necessarily equal

[0073] (5) The worker submits the task information that he is willing to do to the platform

[0074] (6) The platform recruits workers. The method for the platform to recruit workers is as follows: For a certain task, if there is at least 1 worker among the candidates belonging to U A , then only 1 worker from U A should be recruited for this task, and the remaining κ0 - 1 workers are recruited from the set U B ; but if all the workers who declare this task in this round do not have workers belonging to U A , then only 1 worker is recruited for this task. Where κ0 is the number of workers to be recruited, and the value range is 3 - 20

[0075] (7) The platform notifies the selected workers to do the task. After the workers perceive the data, they submit the data to the platform, and the platform obtains the data

[0076] (8) The platform updates the trust level of the workers: After receiving the data, the platform calculates the trust level of the workers. The factor that determines whether the trust level rises or falls is the matching rate between the worker's data and the reference data, that is, the accuracy rate: a i,j,k . The reference data refers to the data submitted by the worker j whose data received by the platform belongs to the set U A , and d i,j,k is the data submitted by the worker belonging to the set U B ; Denote a i,j,k as the accuracy rate, and τ a as the accuracy threshold for accepting this data. When the accuracy rate reaches the threshold set by the platform, it is considered that the data collection behavior of this worker is normal, and his trust level is increased, otherwise it is decreased. The formulas for the increase and decrease of the trust level are as shown in Equation 8

[0077]

[0078] denotes the trust level of worker i for task k in the j - th round, θ T denotes the parameter denotes the maximum trust level

[0079] (8) The platform pays the reward to the worker

[0080] (8.1) The platform determines which stage the worker belongs to and then gives different rewards accordingly. The platform defines two stages for the worker. One is the investment stage, denoted as the DT stage, which means that a certain amount of reward needs to be invested first to make the worker form a path dependence on selecting non-hot tasks. The other is the harvesting stage, denoted as the DM stage, which means that after the platform's initial investment makes the worker form a path dependence on non-hot tasks, the reward probability can be reduced at this time. Since the worker has formed a path dependence and has a lag effect on obtaining rewards, the worker will still select non-hot tasks, so that the investment can be recovered with a lower reward probability. The worker initially belongs to the DT stage. When the worker meets the following two requirements, it is determined that the worker enters the DM stage and the cost is recovered in this stage; otherwise, the worker remains in the DT stage. The first point is: The worker's deviation estimate for α disp is where γ represents the deviation coefficient; when , it can be considered that the worker has formed an estimated deviation for the small-probability event (obtaining an extra reward), which is calculated according to Equation 12 or Equation 16 respectively according to the stage where the worker is located. The second point is that when , it can be considered that the worker has reached the degree of path dependence required by the platform for participating in the non-hot task T np . The path dependence degree evaluation function is as follows:

[0081]

[0082]

[0083] where represents the set of all completed T np tasks by worker i in round j; τ D is the worker's path dependence degree threshold, which is set by the platform; is calculated as shown in formula (3).

[0084] (8.2) If the worker belongs to the DT stage: In the DT stage, the idea of the invention method is to set a higher α r by the platform to increase the actual reward of the worker for completing T np in the DT stage, and then increase the expected utility of completing T np , so that the worker forms a path dependence and enters the DM stage to achieve cost recovery;

[0085] The platform sets the actual winning probability α DT . α DT is higher than the publicly announced probability α disp , so that the posterior probability of the worker's actual winning Always higher than its estimated probability Under this setting, workers will continuously increase their expected probability of obtaining the additional reward for T np As a result, the expected utility of workers for completing T np also continuously rises accordingly. The rising expected utility encourages workers to participate in T np tasks more actively, strengthening the path dependence on participating in T np ; at the same time, it also increases the task completion rate of the system and reduces the proportion of T np .

[0086] Therefore, the estimated probability of workers is a function that is positively correlated with the number of executions of T np . And due to the law of diminishing marginal returns, as continues to increase, its rising speed continues to decline, and finally approaches the true probability α DT infinitely. Therefore, it is defined as as follows:

[0087]

[0088] The actual reward given by the platform to workers is

[0089]

[0090] (8.3) If a worker belongs to the DM stage and enters the DM stage, it means that they have completed the cultivation of path dependence and probability estimation deviation. When a worker has formed a path dependence on executing T np , they are more likely to be affected by path dependence and choose T np , and the weight of profit considered in task selection will relatively decline. Therefore, due to the decreased influence of profit on task selection, the platform keeps α disp unchanged at this stage and reduces the actual winning probability of workers, adjusting it from α DT to α DM (α DT >α disp >α DM ) to reduce costs.

[0091] At the same time, due to the existence of information asymmetry, workers cannot directly know the behavior of the platform to lower the probability. Instead, they need to continuously correct their expected probability through the fact that the actual number of winning awards becomes less in the future, making it gradually approach the true probability α DM . Thanks to the training in the DT stage, the starting point for workers to correct their expected probability in the DM stage is the one finally formed in the DT stage And because ​is a relatively high value, and workers need more time to correct the expected probability closer to the true probability. During this process, the platform can utilize the relatively high expected probability of workers to keep them still maintaining the expectation for T np with high returns, so as to harvest higher-value information with less cost and achieve the purpose of improving the platform's utility.

[0092] The calculation of the expected probability of workers in the DM stage is set as follows:

[0093]

[0094] St: α DM <α DT + [π(α disp ) - α disp

[0095] At the beginning, due to path dependence and inertial thinking, workers are more inclined to use the expected probability cultivated in the DT stage as the expected probability in the DM stage. Therefore, being closer to As the actual probability of the platform in the DM stage decreases, the posterior probability of workers winning awards constantly decreases, and the impact on workers gradually increases, and finally approaches infinitely close to the posterior probability And because of the platform's control Therefore, it will approach infinitely close to α DM ;

[0096]

[0097] The actual reward given by the platform to workers is

[0098]

[0099] (9) If is greater than 0, then start from step (4) to the loop of step (9), and each loop is a round of data collection.

[0100] In the crowd-sourcing sensing network, when it is necessary to obtain the real data of a certain place and requires relatively high data quality, such as in the application scenarios of sensing data such as temperature and humidity, air quality, traffic conditions, etc., and upload the sensed data to the platform for processing. In the tasks released by the platform, participants apply to participate in the tasks, and different data are provided by workers. Applying the method of the present invention in the above data collection can improve data quality, task completion rate, and at a relatively low cost.

[0101] The experimental results of the inventive method are given below.

[0102] ​Figure 1 The update of the worker trust score given by the method of the present invention as the number of task execution rounds increases. It can be seen from the experimental results that the method of the present invention enables the trust score of workers to change with the number of rounds. The trust of trustworthy workers gradually increases, while the trust of malicious workers gradually decreases. This shows that the method of the present invention can identify the trust of workers, so as to guide the recruitment of trustworthy workers when selecting workers to improve data quality.

[0103] Figure 2 The distribution diagrams of the tasks completed by the benchmark method, the benchmark and trustworthy method, the method that only depends on the path without the trustworthy method, and the method of the present invention are given. It can be seen that the task distribution completed by the method of the present invention is the widest, the most uniform, and the task completion rate is high.

[0104] Figure 3 The data quality of the method of the present invention under different worker ratios is given. The meanings of the letters N:M:G in the figure represent: normal workers, malicious workers, and initial trustworthy workers. N:M:G represents the proportion of the three types of worker points. When the composition of different types of workers N:M:G = 240:50:10, the average data accuracy rate of the method of the present invention is 0.96, which is 10.42%, 9.38%, and 10.42% higher than the benchmark method, the benchmark and trustworthy method, and the method that only depends on the path without the trustworthy method respectively. When N:M:G = 190:100:10, the average data accuracy rate of the method of the present invention is 0.94, which is 24.63%, 20.21%, and 21.28% higher than the benchmark method, the benchmark and trustworthy method, and the method that only depends on the path without the trustworthy method respectively. When N:M:G = 140:150:10, the average data accuracy rate of the method of the present invention is 0.89, which is 32.58%, 24.72%, and 29.21% higher than the benchmark method, the benchmark and trustworthy method, and the method that only depends on the path without the trustworthy method respectively. When N:M:G = 90:200:10, the average data accuracy rate of the method of the present invention is 0.93, which is 50.54%, 41.94%, and 44.09% higher than the benchmark method, the benchmark and trustworthy method, and the method that only depends on the path without the trustworthy method respectively.

[0105] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, refer to the description in the method part.

[0106] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A reliable path - dependence - based method for improving data quality and completion rate, characterized in that, Including the following steps: Construct a set of workers. The initial set of workers includes initial trusted workers and workers with undetermined trust levels; The platform releases tasks to be executed. The tasks to be executed are divided into hot tasks and non-hot tasks; Formulate a reward strategy based on the task completion situation of workers and the types of tasks executed; Workers select intended tasks according to their own situations and submit them to the task allocation platform; The platform selects workers and notifies the selected workers to perform tasks. After workers perceive data, they submit the data to the platform, and the platform obtains the data; The platform updates the trust levels of workers according to the task completion situations of workers and pays remuneration.

2. A reliable path-dependence-based method for improving data quality and completion rate according to claim 1, characterized in that The initial set of workers is U, and the initial set of trusted workers is U A , with a trust level of 1. The workers with undetermined trust levels are U B , with an initial trust level of 0.

5. The trust level is between [0, 1], where 0 indicates untrusted.

3. A reliable path-dependence-based method for improving data quality and completion rate according to claim 1, characterized in that The platform releases tasks to be executed Task t i contains the following information: where {x i , y i} represents the coordinates of the task, represents the budget; the platform classifies tasks into a set of hot tasks and a set of non - hot tasks according to the historical completion situation of the tasks: Set of non - hot tasks: Tasks to be executed Hot tasks refer to tasks with an execution intention higher than the preset value and a completion probability higher than the preset value; non - hot tasks refer to tasks with an execution intention lower than the preset value and a completion probability lower than the preset value.

4. A reliable path - dependence - based method for improving data quality and completion rate according to claim 1, characterized in that, A reward strategy is formulated according to the task completion situation of workers and the types of tasks they perform. Specifically: The platform formulates the payment to workers As shown in Equation (1), represents the payment to worker i for completing task k in the j - th round of data collection announced by the platform; To encourage workers to participate in more tasks in set T np Therefore, if the worker participates in a task in T p in this round, the payment they receive is only the fixed marked price B; While if the worker participates in a task in T np in this round, then on the basis of the fixed marked price B, there is a probability of α disp to obtain an additional reward. α disp is the winning probability publicly shown by the platform to workers; d i,k represents the distance between worker i and task k. When d i,k = 0, the moving cost of the worker should be 0; δ p represents the unit - distance moving cost coefficient; In addition, the actual probability for a worker set by the platform to obtain an additional reward: α DT is not publicly announced by the platform; The specific tasks selected by the worker include: the worker traverses all tasks within his own movement range, calculates the corresponding R for each task in this round i,j,k (·), and then selects max(R i,j,k (·)) as the task willing to do in the j-th round; The definition of R i,j,k (·) is as follows: Among them, represents the expected utility of the worker for this task in this round; is calculated as shown in formula (3); d i,k represents the distance between worker i and task k; D i,j represents the degree of path dependence of the worker in this round. The deeper the worker's path dependence on , the more inclined the worker is to choose task; D i,j is calculated as shown in formula (14); ω U , ω d , ω D are parameters representing the parameters of expected utility, distance, and degree of path dependence; The cost of the worker is shown in Equation (4). The cost of the worker participating in the task consists of two parts, namely the fixed data acquisition cost A and the movement cost. d i,k represents the distance between the worker and the task; when d i,k = 0, the movement cost of the worker should be 0. Therefore, the movement cost is set to δ c represents the movement cost coefficient per unit distance, represents the reward estimated by the worker himself, represents the reward probability estimated by the worker himself. The probability of obtaining an additional reward is related to α announced by the platform disp .

5. A reliable path - dependence - based method for improving data quality and completion rate according to claim 1, characterized in that, The method for the platform to select workers is as follows: For any task, if there is at least 1 worker among the candidates who belongs to the initial set of trusted workers, only 1 worker from the set of trusted workers is recruited, and the remaining κ0 - 1 workers are recruited from the set U B ; if none of the workers who declared this task in this round belong to U A , then only 1 worker is recruited for this task; where κ0 is the number of workers to be recruited.

6. A reliable path - dependent method for improving data quality and completion rate according to claim 1, characterized in that, The platform updates the trust level of workers based on their task completion. Specifically, the factor determining whether the trust level increases or decreases is the matching rate between the worker data and the baseline data, where the baseline data refers to the data received by the platform that belongs to the set U A submitted by worker j, and d i,j,k is the data submitted by a worker belonging to the set U B ; let a i,j,k represent the accuracy rate, and τ a represent the accuracy threshold for accepting this data. When the accuracy rate reaches the threshold set by the platform, it is considered that the data collection behavior of the current worker for this data is normal, and their trust level is increased; otherwise, it is decreased. The formulas for increasing and decreasing the trust level are as shown in Equation (8): represents the trust level of worker i in task k at the j - th round, θ T represents a parameter, represents the maximum trust level.

7. A reliable path-dependence-based method for improving data quality and completion rate according to claim 4, characterized in that, The specific steps for paying remuneration are as follows: The platform defines two stages for workers. One is the investment stage, denoted as the DT stage, and the other is the harvesting stage, denoted as the DM stage. Workers initially belong to the DT stage. When a worker meets the following requirements, it is determined that the worker enters the DM stage and recovers the cost in this stage; otherwise, the worker remains in the DT stage. The worker's deviation prediction for α disp is where γ represents the deviation coefficient. When , it is considered that the worker has formed a prediction deviation for low-probability events. is calculated according to Equation (12) or Equation (16) respectively according to the stage where the worker is located. When , it is considered that the worker has reached the degree of path dependence required by the platform for participating in the non-hot task T np . The path dependence degree evaluation function is as follows: Among them, represents the set of all completed T np tasks of worker i at round j; τ D is the threshold of the worker's path dependence degree, which is set by the platform; is calculated as shown in formula (3); If the worker belongs to the DT stage: In the DT stage, the idea of the invention method is to set a relatively high α through the platform r , to increase the actual reward of the worker for completing T np in the DT stage, and then to enhance the expected utility of completing T np , so that the worker forms path dependence and proceeds to the DM stage to achieve cost recovery; The platform sets the actual winning probability α DT ; α DT is higher than the publicly announced probability α disp , so that the posterior probability of the worker's actual winning is always higher than his estimated probability Under this setting, the worker will continuously increase the expected probability of obtaining the T np extra reward, and the worker's expected utility for completing T np also continuously increases accordingly; the increase in expected utility encourages the worker to participate in T more actively np tasks, strengthening the path dependence on participating in T np ; at the same time, it also increases the task completion rate of the system and reduces the proportion of T np ; Therefore, the estimated probability of the worker is a function that is positively correlated with T np the number of executions ; and due to the diminishing marginal effect, as continues to increase, its rising speed continues to decline and finally approaches the true probability α DT infinitely; therefore, it is defined as as follows: The actual reward given to workers by the platform is If the worker belongs to the DM stage and enters the DM stage, it indicates that the cultivation of path dependence and probability estimation deviation has been completed. When the worker has developed a path dependence on performing T np The platform maintains α disp unchanged at this stage and reduces the actual winning probability of the worker from α DT to α DM , where α DT >α disp >α DM , thereby reducing costs; The calculation of the expected probability for workers in the DM stage is set as follows: St:α DM <α DT +[π(α disp )-α disp ​ At the beginning, due to path dependence and inertial thinking, workers will be more inclined to use the expected probability cultivated in the DT stage as the expected probability in the DM stage. Therefore, closer to As the actual probability of the platform in the DM stage decreases, the posterior probability of workers winning the award continually decreases, and the impact on workers gradually increases, and finally approaches the posterior probability And because of platform control therefore will approach α infinitely DM ; The actual reward given to workers by the platform is

Citation Information

Patent Citations

  • Quality guarantee method for crowd sensing data based on endowment effect and reference dependence

    CN114866552A

  • Crowd sensing excitation method, system and equipment based on path dependence theory

    CN114926088A

  • Mobile crowd sensing system quality guarantee method based on sheep flock effect, equipment and medium

    CN115271629A