Resource-optimized dynamic crowd sensing truth value discovery method

By using reinforcement learning to dynamically publish the Gold Standard Question and iterative weighted aggregation, the problems of rigidity and insufficient anti-interference ability of the GSQ strategy in crowd perception are solved, enabling continuous estimation of user capabilities and efficient inference of truth values, thus improving the stability and accuracy of the crowd perception system.

CN121996914APending Publication Date: 2026-05-08KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
KUNMING UNIV OF SCI & TECH
Filing Date
2026-01-07
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In existing crowd sensing technologies, the GSQ distribution strategy lacks dynamic adaptability, leading to resource waste or insufficient calibration. Traditional truth discovery methods have weak anti-interference capabilities, affecting the accuracy of truth inference.

Method used

By using reinforcement learning to dynamically distribute the Gold Standard Question (GSQ), and combining iterative weighted aggregation methods to optimize user capability estimation and truth inference, a reinforcement learning model is designed to dynamically adjust the GSQ distribution strategy and user weight allocation to prevent malicious attacks.

Benefits of technology

It enables continuous estimation of user capabilities and accurate inference of truth values ​​in high-noise and malicious attack environments, optimizes resource utilization efficiency, balances calibration accuracy and cost, and improves the robustness and accuracy of truth discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121996914A_ABST
    Figure CN121996914A_ABST
Patent Text Reader

Abstract

The invention relates to a resource-optimized dynamic crowd sensing truth value discovery method, and belongs to the field of crowd sensing and crowdsourcing data management. The method comprises the following steps: obtaining initial capability estimation values of users through initial golden standard problem tasks, sorting the initial capability estimation values, and distributing common perception tasks for the users; and performing iterative calculation on perception data of all users participating in the same common perception task to obtain final truth value estimation of the current common perception task, thereby calculating a temporary capability index of each user, constructing a state vector and inputting the state vector into a reinforcement learning model to obtain a decision of a golden standard problem calibration task, and finally obtaining a standard problem calibration task. A gold standard problem calibration task is allocated to the user; and generating a reward signal based on the task release cost to update parameters of the reinforcement learning model, and optimizing generation of a decision. The objective of the invention is to solve the technical problems that in the prior art, a task allocation strategy cannot be dynamically adjusted, and resource waste or insufficient calibration precision is caused by a fixed-frequency GSQ publishing mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a dynamic collective intelligence perception truth discovery method for resource optimization, belonging to the fields of collective intelligence perception and crowdsourced data management. Background Technology

[0002] The rapid development of internet and mobile computing technologies has driven the widespread application of crowdsensing and crowdsourcing models in fields such as environmental monitoring, traffic management, business decision-making, and data annotation. Crowdsensing relies on a large number of users collecting and submitting sensing data through smart devices to complete various tasks. However, due to varying user capabilities, interference from ambient noise, and the potential for malicious attacks, the collected data suffers from low quality and poor reliability, directly affecting the accuracy of subsequent data analysis and decision-making.

[0003] To improve data quality, truth discovery techniques are widely used to estimate the true value of a task from noisy answers from multiple users. Well-known truth discovery methods include simple averaging and weighted averaging. Furthermore, to calibrate user capabilities and reduce system errors, Panagiotis G. Ipeirotis and Evgeniy Gabrilovich (2014. Quizz: Targeted crowdsourcing with a billion (potential) users. In Proceedings of the 23rd International Conference on World Wide Web (WWW'14). ACM, 143–154) proposed the Gold Standard Questions (GSQ) as a benchmark task. The true value of the GSQ is known, and user reliability is assessed by comparing user answers with the true value. However, existing techniques have significant shortcomings in both GSQ distribution strategies and data aggregation algorithms, making them difficult to meet the needs of complex scenarios.

[0004] On the one hand, the existing GSQ distribution lacks dynamic adaptability, mostly adopting a fixed frequency or random distribution mode, without dynamically adjusting in combination with user status and system resource constraints. This rigid strategy is prone to two types of problems: either excessive GSQ distribution increases system costs, or insufficient GSQ distribution fails to track changes in user capabilities in a timely manner, leading to the failure of user capability estimation, which in turn amplifies the true value inference bias and makes it difficult to balance the relationship between calibration accuracy and resource costs.

[0005] On the other hand, traditional truth-finding aggregation algorithms are weak in their ability to withstand interference, relying heavily on static weight allocation or simple outlier filtering mechanisms, which cannot effectively cope with high-noise environments and malicious user attacks. When faced with noisy data such as equipment measurement errors, static weights are unable to dynamically mitigate the impact of low-quality data, leading to a significant decrease in the accuracy of truth inference and insufficient robustness.

[0006] This invention, with the background of data quality control for crowd-aware sensing, proposes a resource-optimized dynamic truth discovery method for crowd-aware sensing. By making intelligent decisions on "when and to whom to issue GSQ" and combining it with an iterative weighted aggregation method based on user capabilities and data errors, it achieves continuous estimation of user capabilities and accurate inference of truth values ​​in high-noise and malicious attack environments. This provides a new technology for solving the problems of existing methods in terms of dynamic adaptability, resource efficiency, and resistance to malicious attacks. Summary of the Invention

[0007] The purpose of this invention is to provide a resource-optimized dynamic swarm intelligence truth discovery method, which aims to solve the problems in the prior art where the task allocation strategy fails to be dynamically adjusted, resulting in excessive or insufficient GSQ releases and an inability to adapt to changes in user capabilities in real time; existing data aggregation methods have poor anti-interference capabilities when facing malicious users and high-noise data, affecting the accuracy of truth inference; and fixed-frequency GSQ release methods lead to resource waste or insufficient calibration accuracy, making it difficult to balance system cost and the accuracy of task results.

[0008] To achieve the above objectives, the technical solution of this invention is: a resource-optimized dynamic swarm intelligence truth discovery method. This method uses reinforcement learning to dynamically distribute the gold standard problem to optimize calibration task distribution, and combines robust aggregation to achieve efficient truth inference, thereby improving the inference accuracy of truth discovery and optimizing resource utilization efficiency. It includes the following steps: Step 1: Calculate the user's initial ability estimate based on the error between the user's actual result of completing the initial gold standard problem task and the known true value of the initial gold standard problem task. Step 2: Sort each user's initial capability estimate from highest to lowest, and assign a common perception task to each user according to the sorting result; wherein, the common perception task is completed by multiple users. Step 3: Iteratively calculate the perception data of all users participating in the same common perception task to obtain the final truth estimate of the current common perception task, and calculate the temporary capability index of each user based on the deviation between each user's perception data and the final truth estimate. Step 4: Construct a state vector based on the user's initial ability estimate and temporary ability index, and input the state vector into the reinforcement learning model to obtain the decision of the gold standard problem calibration task output by the reinforcement learning model, so as to assign the gold standard problem calibration task to the user based on the decision result; wherein, the reinforcement learning model includes a state vector, an action space and a reward function; Step 5: Generate a reward signal based on the user's capability estimate and the task release cost, and use the reward signal to update the parameters of the reinforcement learning model in order to optimize the generation of decisions based on the updated parameters.

[0009] Optionally, the expression for the user's initial capability estimate is: in, For users The initial capability estimate, The error between the user's measured value and the actual value from the calibration task. The preset upper limit of error, This is the lower limit threshold for the estimated user capability value.

[0010] Optionally, Step 3 specifically includes: Step 3.1: Based on the users' initial ability estimates, obtain the initial weights for each user participating in the same general perception task, and normalize them. The expression is: in, For users The initial weights, The number of users participating in the same common perception task in this round. This is a user index variable that participates in the same general perception task in this round, with a value range of 1 to... This is used to refer to each user individually; Step 3.2: Calculate the weighted average based on the initial weights as the first aggregation result. The expression is: in, This is the initial aggregate estimate. For users The measured value; Step 3.3: Proceed In the next iteration, the absolute error between each user's measurement and the current aggregation result. for: in, For the first The aggregated estimate of the next iteration is weighted according to the magnitude of the error during the iteration process. The error weights are calculated as follows: in, For users In the The error weights obtained in the next iteration It is a constant; Step 3.4: Combine the error weights with the user capability estimates to obtain the comprehensive weight. The formula for calculating the comprehensive weight is: in, For users In the The comprehensive weights obtained in the nth iteration are normalized to obtain the nth iteration. Weight of +1 iteration The expression is: Step 3.5: Calculate the first... The result of +1 iteration yields a new aggregation result. The expression is: Step 3.6: Stop the iteration when the maximum relative change of the weights is less than the preset iteration convergence threshold. The expression is: in, This is the threshold for iterative convergence; Step 3.7: After the user completes the basic perception task, a temporary capability index is obtained based on the error between the user's measurement value and the final aggregated result. The expression is as follows: in, For users Temporary ability value, This represents the error between the user's measured value and the final aggregated result.

[0011] Optionally, the reinforcement learning model is specifically: State vector S i= [ability, interval, variance, recent_mae], where ability is the user's current ability estimate, reflecting the quality level of the user's task completion; interval is the number of ordinary perception tasks the user has completed since the last completion of the gold standard problem task; variance is the standard deviation of the user's ability estimates over the last n times, reflecting the uncertainty or volatility of the ability estimates; and recent_mae is the average absolute error of the user in the last n ordinary perception tasks, reflecting the error between the user's perception data and the aggregated true value over the last n times. Action space A = {issue_GSQ, not_issue}, where issue_GSQ represents the gold standard issue to be published for the user, used to calibrate the user's capabilities; not_issue represents not publishing a gold standard issue for the user, and the user will continue to perform the ordinary perception tasks published by the platform. The reward function is reward = (base_reward + GSQ_total + stability_reward), where base_reward reflects whether the Gold Standard Question (GSQ) release decision improves the accuracy of user capability estimation; GSQ_total controls the effectiveness and frequency of GSQ use to ensure that GSQ releases are valuable; and stability_reward encourages the generation of stable and reliable capability estimates.

[0012] Optionally, the expression for base_reward is: in, The threshold is used to determine whether the adjustment range is within a predefined range. It is a natural constant. Estimate new capabilities and old capacity estimates The absolute difference is used to measure the magnitude of the adjustment; This is a positive reward coefficient used to adjust the reward intensity. This is a negative penalty coefficient used to adjust the intensity of the penalty. and The error attenuation coefficient is used to control the sensitivity of rewards and penalties as the adjustment range changes.

[0013] Optionally, the expression for GSQ_total is: in, The expression for the GSQ effect bonus is: in, This represents the effective reward coefficient for GSQ. This represents the GSQ invalidity penalty coefficient. and The intensity of rewards and penalties used to control the effectiveness of GSQ calibration; To effectively calibrate the error threshold, Invalid calibration error threshold, and The error threshold used to determine whether a GSQ is valid or invalid. The absolute difference between the new capability estimate and the old capability estimate; in, The expression for GSQ frequency reward is: in, To accumulate the number of GSQs published for users, , The threshold for GSQ frequency segmentation; To facilitate the early use of reward coefficients, To avoid overuse of the penalty coefficient, and The intensity of rewards and penalties used to control the frequency of GSQ usage; This is the attenuation coefficient for excessive penalties.

[0014] Optionally, the expression for stability_reward is: in, The variance of the capability estimate is used to measure the stability of the user's capability. This serves as the base coefficient for stability rewards, used to control the upper limit of rewards for stability capabilities. This is the stability decay coefficient, used to control the variance of reward as a function of ability estimation. Increased decay rate.

[0015] The beneficial effects of this invention are: 1. Existing GSQ (Gauge Scale) distribution methods, using either a fixed frequency or random distribution pattern, either increase system costs due to over-distribution or fail to track user capability changes in a timely manner due to insufficient distribution, leading to calibration failure. This invention achieves dynamic decision-making for GSQ distribution through a reinforcement learning agent: based on user state vectors, the agent accurately determines whether GSQ distribution is necessary and which users should receive GSQs first. Users with significant capability fluctuations, long-term lack of calibration, or high recent task errors are given priority for GSQ calibration; users with stable capabilities and small recent errors are temporarily exempt from GSQ distribution to conserve resources. This strategy ensures the accuracy of user capability calibration while avoiding meaningless GSQ resource consumption, perfectly balancing calibration effectiveness and cost control, and solving the inherent defects of over-consumption or insufficient calibration in fixed or random distribution modes.

[0016] 2. Traditional aggregation methods rely on static weight allocation or simple outlier filtering, which struggles to handle high-noise data and malicious users, leading to decreased accuracy in truth inference. This invention strengthens anti-interference capabilities through two major designs: First, it employs an iterative weighted aggregation method based on user ability and data error. Initial weights are assigned based on the user's current ability value, and then through multiple iterations, the weights are dynamically adjusted according to the deviation between the user's answer and the temporary aggregation result. This ensures that users with smaller errors and higher abilities receive greater overall weights, proactively identifying and reducing the interference of low-quality data on the truth. Second, it designs a confusion mechanism between ordinary perception tasks and GSQ tasks. The two types of tasks are completely identical in their interaction format (no data sharing between users required) and reward mechanism (rewards are issued based on task performance). Users cannot distinguish between types based on task characteristics, completely avoiding malicious users' attacks that deliberately circumvent GSQ calibration and submit only false data for ordinary perception tasks. This improves the robustness of truth inference in high-noise, multi-interference scenarios, ensuring the reliability of the final truth result. Attached Figure Description

[0017] Figure 1 This is a flowchart of the present invention; Figure 2 This is a comparison chart of the mean absolute error of the task under different strategies according to the present invention; Figure 3 This is a comparison chart of the usage of the gold standard problem under different strategies of this invention; Figure 4 This is a comparison chart showing the variation of the mean absolute error of the present invention with the number of users under different aggregation algorithms. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] Example 1: As Figure 1As shown, a resource-optimized dynamic crowd-aware truth-finding method includes the following steps: Step 1: Calculate the user's initial ability estimate based on the error between the user's actual result of completing the initial gold standard problem task and the known true value of the initial gold standard problem task. Optionally, the expression for the user's initial capability estimate is: in, For users The initial capability estimate, The error between the user's measured value and the actual value from the calibration task. The preset upper limit of error, The lower limit threshold for the user capability estimate is set when the error exceeds... hour, Limited to the minimum value The smaller the error between the user's measured value and the final aggregated result, the greater the user's ability value.

[0020] Step 2: Sort each user's initial capability estimate from highest to lowest, and assign a common perception task to each user according to the sorting result; wherein, the common perception task is completed by multiple users. Optionally, the user set is sorted in descending order of the initial capability estimates to obtain: in, for A set of users in descending order; And the sorting satisfies: in, , and Representing users respectively ,user With users Capacity estimates.

[0021] Step 3: Iteratively calculate the perception data of all users participating in the same common perception task to obtain the final truth estimate of the current common perception task, and calculate the temporary capability index of each user based on the deviation between each user's perception data and the final truth estimate. It should be understood that the perceived data is the data submitted by the user for the assigned general perception task; Optionally, this embodiment employs an iterative weighted aggregation method based on user capabilities and data errors for truth discovery and reward allocation in ordinary perception tasks, specifically as follows: Step 3.1: Based on the users' initial ability estimates, obtain the initial weights for each user participating in the same general perception task, and normalize them. The expression is: in, For users The initial weights, The number of users participating in the same common perception task in this round. This is a user index variable that participates in the same general perception task in this round, with a value range of 1 to... This is used to refer to each user individually; Step 3.2: Calculate the weighted average based on the initial weights as the first aggregation result. The expression is: in, This is the initial aggregate estimate. For users The measured value; It is understood that the initial aggregated estimate is a temporary truth inference based on the user's initial ability weights and the first submitted answer. It is not the final truth value of the task and needs to be converged to a stable value through subsequent iterative optimization. Step 3.3: Proceed In the next iteration, the absolute error between each user's measurement and the current aggregation result. for: in, For the first The aggregated estimate of the next iteration is weighted according to the magnitude of the error during the iteration process. The smaller the error, the larger the weight. The error weight is calculated as follows: in, For users In the The error weights obtained in the next iteration It is a constant, a local minimum value, used to avoid division by zero in the denominator. In this embodiment, Pick ; Step 3.4: Combine the error weights with the user's ability estimate to obtain the comprehensive weight. Users with higher abilities and smaller errors have larger weights. The formula for calculating the comprehensive weight is: in, For users In the The comprehensive weights obtained in the nth iteration are normalized to obtain the nth iteration. Weight of +1 iteration The expression is: Step 3.5: Calculate the first... The result of +1 iteration yields a new aggregation result. The expression is: Step 3.6: Stop the iteration when the maximum relative change of the weights is less than the preset iteration convergence threshold. The expression is: in, This is the iterative convergence threshold, used to determine whether the weight changes are small enough and whether the aggregation result is stable. Step 3.7: After the user completes the basic perception task, a temporary capability index is obtained based on the error between the user's measurement value and the final aggregated result. The expression is as follows: in, For users Temporary ability value, The error between the user's measured value and the final aggregated result; when the error exceeds... hour, Limited to the minimum value The smaller the error between the user's measured value and the final aggregated result, the greater the user's ability value.

[0022] It is understandable that the purpose of introducing temporary capability values ​​in this embodiment is to calculate the fluctuation state of user capability values ​​so as to input them into the reinforcement learning model and optimize the calibration task release strategy. In the subsequent execution of ordinary perception tasks, the initial weights are still calculated from the capability values ​​obtained after executing the gold standard problem task.

[0023] Optionally, this embodiment also distributes pre-deposited rewards to corresponding users based on their contributions to the task, completing the incentive loop. The user reward distribution formula is as follows: in, For users Rewards for completing the task The total reward for this task will be paid by the user who initiated the task. For users during the iterative aggregation calculation of the actual value of the task The final weight is determined by the user's ability. The higher the weight (the more capable the user is and the smaller the deviation between their answer and the final aggregated result), the more reward they will receive.

[0024] Understandably, Step 3 of this embodiment employs an iterative weighted aggregation method based on user ability and data error. In the aggregation of the crowd-sensing task, initial weights are first obtained based on user ability estimates. Then, a weighted average result is calculated through multiple iterations. After each iteration, the weights are readjusted based on the deviation between the current aggregation result and the user's answer, so that users with better performance receive higher weights, while the weights of users with poorer performance gradually decrease. Furthermore, the error weights are combined with the user ability values ​​to obtain a comprehensive weight; users with higher abilities and smaller errors receive higher weights. After multiple iterations, the aggregation finally converges to a stable result. This iterative weighted aggregation method can effectively identify and reduce the impact of malicious or low-ability users on the final result, improving the accuracy and robustness of the crowd-sensing task aggregation.

[0025] Step 4: Construct a state vector based on the user's initial ability estimate and temporary ability index, and input the state vector into the reinforcement learning model to obtain the decision of the gold standard problem calibration task output by the reinforcement learning model, so as to assign the gold standard problem calibration task to the user based on the decision result; wherein, the reinforcement learning model includes a state vector, an action space and a reward function; It is important to understand that each gold standard problem calibration task is completed by a single user. For the ability calibration and reward allocation of the gold standard problem calibration task: the error is calculated based on the user's submitted answer and the known truth value of the task, and the user's current ability estimate is updated based on this error. ; Furthermore, rewards are distributed to users based on their performance in completing calibration tasks, incentivizing participation and ensuring data quality. Since calibration tasks are published by the platform, and the data requester is not involved in the process, the task rewards are distributed by the platform. When the platform publishes ordinary sensing tasks to data requesters, it will distribute rewards according to a preset ratio. (GSQ Reward Commission Rate) This refers to the percentage of the total reward amount charged to data requesters for ordinary perception tasks. This is to reward users for performing calibration tasks. Since users do not need to communicate or share data with other users during the performance of both normal perception tasks and calibration tasks, and both tasks are rewarded, users cannot distinguish between normal perception tasks and calibration tasks, and therefore cannot carry out malicious actions.

[0026] Optionally, the reinforcement learning model is specifically: State vector S i= [ability, interval, variance, recent_mae], where ability is the user's current ability estimate, reflecting the quality level of the user's task completion; interval is the number of ordinary perceptual tasks the user has completed since the last completion of the gold standard problem task; and variance is the standard deviation of the user's most recent n ability estimates. , used to reflect the uncertainty or volatility of capability estimation; recent_mae is the average absolute error value of the user in the most recent n ordinary perception tasks, used to reflect the error between the user's perception data and the aggregated true value in the most recent n tasks; Specifically, standard deviation The calculation formula is: in, The average of the user's most recent n temporary ability values; Action space A = {issue_GSQ, not_issue}, where issue_GSQ represents the gold standard issue to be published for the user, used to calibrate the user's capabilities; not_issue represents not publishing a gold standard issue for the user, and the user will continue to perform the ordinary perception tasks published by the platform. The reward function is reward = (base_reward + GSQ_total + stability_reward), where base_reward reflects whether the Gold Standard Question (GSQ) release decision improves the accuracy of user capability estimation; GSQ_total controls the effectiveness and frequency of GSQ use to ensure that GSQ releases are valuable and not overused; and stability_reward encourages the generation of stable and reliable capability estimates and avoids drastic fluctuations in capability values.

[0027] Optionally, the expression for base_reward is: in, The threshold is used to determine whether the adjustment range is within a predefined range. It is a natural constant. Estimate new capabilities and old capacity estimates The absolute difference is used to measure the magnitude of the adjustment; This is a positive reward coefficient used to adjust the reward intensity. This is a negative penalty coefficient used to adjust the intensity of the penalty. and The error attenuation coefficient is used to control the sensitivity of rewards and penalties as the adjustment range changes.

[0028] Optionally, the expression for GSQ_total is: in, The expression for the GSQ effect bonus is: in, This represents the effective reward coefficient for GSQ. This represents the GSQ invalidity penalty coefficient. and The intensity of rewards and penalties used to control the effectiveness of GSQ calibration; To effectively calibrate the error threshold, Invalid calibration error threshold, and The error threshold used to determine whether a GSQ is valid or invalid. The absolute difference between the new capability estimate and the old capability estimate; Among these measures, rewards and penalties are based on whether the error is improved or worsened after GSQ implementation. The expression for GSQ frequency reward is: in, To accumulate the number of GSQs published for users, , The GSQ frequency segmentation threshold is used to divide the critical number of times GSQ is used in the early, middle and excessive use stages. To facilitate the early use of reward coefficients, To avoid overuse of the penalty coefficient, and The intensity of rewards and penalties used to control the frequency of GSQ usage; This is the over-penalty decay coefficient, used to control the gradient of the penalty as the number of uses increases when overuse occurs.

[0029] Optionally, stability_reward is a capability stability reward that encourages the stability of capability estimates. The smaller the variance of the capability value, the higher the reward. The expression for stability_reward is: in, The variance of the capability estimate is used to measure the stability of the user's capability. This serves as the base coefficient for stability rewards, used to control the upper limit of rewards for stability capabilities. This is the stability decay coefficient, used to control the variance of reward as a function of ability estimation. Increased decay rate.

[0030] It is understandable that the core role of the reward function introduced in this embodiment is to guide the Q-learning agent to learn the optimal gold standard question release strategy, so that the reinforcement learning agent can adaptively learn when to release the GSQ most effectively in a complex dynamic environment. This balances the efficiency and cost of the system while ensuring the accuracy of user capability estimation, and ultimately achieves continuous optimization of the overall system performance.

[0031] Step 5: Generate a reward signal based on the user's capability estimate and the task release cost, and use the reward signal to update the parameters of the reinforcement learning model in order to optimize the generation of decisions based on the updated parameters.

[0032] Example 2: Based on the technical solution provided in Example 1, this example further illustrates the present invention through a specific implementation example.

[0033] Specifically, taking the real-time temperature monitoring of the city center area by the environmental monitoring department of a certain city as an example, the core roles and task objectives are as follows: Data requester: Environmental monitoring department of a certain city, needs to obtain the actual temperature of the area.

[0034] Users: 100 citizens equipped with temperature sensors (numbered W1-W100) must first complete capability calibration through the Gold Standard Question (GSQ) before they can accept ordinary sensing tasks.

[0035] The Crowd Intelligence Perception Platform is responsible for task distribution, data aggregation, and reward distribution. It integrates a reinforcement learning model (determining the timing and recipients of GSQ rewards) with an iterative weighted aggregation method based on user capabilities and data errors (calculating the true value of the task).

[0036] Key parameters: GSQ Reward Commission Rate: 5% (deducted from the rewards of regular perception tasks and deposited into the GSQ exclusive reward pool); User capability threshold ( ): 0.1 (to avoid data weights becoming invalid due to excessively low capability values); Reasonable upper limit of temperature error ( ): 2℃ (when the measurement error exceeds 2℃, the user capability is calculated as 0.1); Convergence threshold of aggregation algorithm ( ): 2% (Stop iterative calculation when the user weight change is less than 2%) Furthermore, taking the task of monitoring the current temperature in area B initiated by the environmental monitoring department as an example, the implementation process is as follows: S1: Capability Initialization Taking user W1 as an example, a GSQ task (preset true value 25.0℃) for region A is issued to newly registered user W1; W1 submits data of 25.2℃ after on-site measurement; the initial capability value is calculated. With an error of 0.2℃, the calculated capability value is 0.9. Save W1's initial ability value (0.9) and distribute the reward from the GSQ reward pool.

[0037] The environmental monitoring department submitted its task request through the platform: to monitor the current temperature of area B, and transferred the task reward of 500 yuan to the platform.

[0038] S2: Task Assignment When assigning tasks, users are sorted in descending order of their ability values, and tasks are assigned to users with higher ability values ​​first. This ensures that the data obtained from the tasks is of higher quality, and thus the aggregated task results are closer to the actual task values.

[0039] The platform selects the top 5 users with the highest ability scores from the available users, and sorts them in descending order of ability: W1 (0.9), W2 (0.85), W3 (0.8), W4 (0.75), and W5 (0.7), and assigns tasks to these 5 users; S3: Sensing data reception After receiving the task, the user completes the on-site measurement and submits the data: W1 (24.8℃), W2 (25.1℃), W3 (24.7℃), W4 (25.3℃), W5 (24.9℃).

[0040] S4: Data Aggregation and Reward Allocation The core logic of the iterative weighted aggregation method based on user capability and data error is that users with higher capabilities and smaller measurement errors have greater data weights. The steps are as follows: Initial weight allocation: Based on user ability values, W1 has the highest weighting of 0.225, and W5 has the lowest weighting of 0.175. Iterative weight adjustment: Calculate the temporary average temperature of 24.955℃, compare the error between the user's measured value and this temperature to obtain the error weight, and combine it with the user's capability value to obtain the comprehensive weight. Repeat the iteration. Stop iteration: Repeat the adjustment until the weight change is less than 2%, and finally obtain the true value of the task in region B as 24.92℃ (with an error of only 0.02℃ from the actual temperature of 24.9℃).

[0041] Calculate the user's temporary capability value. The temporary capability value is only used for reinforcement learning state analysis and does not update the formal capability value. For example, W5 error is 0.02℃, and the temporary capability value is 0.99; W4 error is 0.38℃, and the temporary capability value is 0.81.

[0042] The platform reported to the environmental monitoring department that the final temperature in area B was 24.92℃; The reward of 500 yuan will be distributed according to the final data weight of the users: W5 has the highest weight (42.8%) and will receive 214 yuan; W1 has a weight of 22.5% and will receive 97 yuan; at the same time, 5% (25 yuan) will be deposited into the GSQ reward pool.

[0043] S5: User Status Tracking Taking user W3 as an example, the platform passes W3's state vector to the reinforcement learning model: Assume that W3 has completed 8 normal perception tasks since the last GSQ, its recent ability values ​​have fluctuated greatly, and its recent error is relatively high; S6: Calibration Decision Generation The reinforcement learning model determined that the benefit of publishing GSQ calibration was higher than not publishing, and the feedback platform assigned a GSQ task to W3.

[0044] S7: Task Assignment and Execution The platform issues a GSQ task (preset truth value 26.5℃) for region C to W3. The task format is the same as that of a normal perception task. W3 submitted a measurement of 26.6℃ to the crowd sensing platform, and the platform's calculation error was 0.1℃. The user's ability value is updated according to the ability value calculation formula, and the new ability value of W3 is calculated to be 0.95. The platform updates the ability value in W3's profile from 0.8 to 0.95. The crowd-sensing platform distributes rewards to user W3 from the GSQ reward pool; S8: Strategy Iteration Optimization Reward signals are generated based on the update effect of user capability estimates and task release costs. The parameters of the reinforcement learning model are then updated using these reward signals to optimize the generation of subsequent calibration decisions.

[0045] By cyclically executing the above processes S1-S8, the dynamic tracking of user capabilities and continuous optimization of truth inference are achieved, ultimately improving the stability and accuracy of the crowd perception system in high-noise and multi-interference scenarios.

[0046] Furthermore, Figure 2 The results show the mean absolute error (MAE) obtained from experiments using reinforcement learning-based GSQ deployment policies, fixed low-frequency GSQ deployment policies, fixed high-frequency GSQ deployment policies, random GSQ deployment policies, and no GSQ policy. Figure 3This paper demonstrates the number of GSQs used in 500 tasks across various reinforcement learning strategies: fixed low-frequency GSQ deployment, fixed high-frequency GSQ deployment, random GSQ deployment, and no GSQ deployment. The experiment constructed a swarm intelligence sensing platform and used temperature data from the real-world IntelLAB Data dataset as the ground truth for each task. The mean absolute error between the aggregated results and the ground truth was calculated as the performance metric. The four GSQ-based strategies all employed the iterative weighted aggregation algorithm fused with the present invention (IWLS) when aggregating user-perceived data. The no-GSQ strategy, lacking initial weights, used the average weighted aggregation method from traditional aggregation methods.

[0047] Experimental results show that, across 500 tasks, the reinforcement learning strategy achieved the lowest MAE value of 1.2 with only 32 GSQ usages. This is significantly better than the fixed high-frequency strategy (GSQ is issued once every 3 tasks, using 167 GSQ usages) at 1.9, the random strategy (GSQ is issued with a 20% probability on each task) at 2.1 (using 101 GSQ usages), and the fixed low-frequency strategy (GSQ is issued once every 10 tasks, using 50 GSQ usages) at 2.3. The MAE of the strategy without GSQ is as high as 4.2, fully validating the necessity of the calibration mechanism. This result demonstrates the core innovative value of the dynamic GSQ strategy: by learning the optimal GSQ release timing and target selection through Q-learning agents, the system can maximize resource utilization efficiency under limited budget constraints. Compared with traditional fixed-frequency or random strategies, the reinforcement learning strategy not only reduces MAE error by 54% but also saves more than 69% of GSQ resources. This provides an economical and efficient data quality control solution for practical crowdsourcing platforms facing complex real-world challenges such as budget constraints, dynamic changes in user capabilities, and noise interference.

[0048] Furthermore, Figure 4This chart compares the mean absolute error (MAE) of different aggregation algorithms under different user numbers, illustrating the impact of various data aggregation algorithms on the MAE under different user numbers. The number of tasks is fixed at 200, and the number of users ranges from 10 to 100. Aggregation algorithms include the method of this invention, the average algorithm, the median algorithm, the CRH algorithm, and the PRTD algorithm. Experimental results show that the MAE of these aggregation algorithms decreases with the increase in the number of users. This is because as the number of users increases, higher-ability users are introduced, and tasks are preferentially assigned to users with higher ability values, resulting in higher data quality from participating users and ultimately a decrease in the MAE. The IWLS algorithm of this invention maintains the lowest MAE across all user numbers (10-100 people), showing an average performance advantage of 38-63% compared to PRTD, CRH, the average algorithm, and the median algorithm. This advantage is particularly pronounced as the number of users increases, with a 44.5% improvement in MAE from 10 to 100 users, demonstrating the significant effect of this invention in improving data quality.

[0049] In summary, this invention first assesses user capabilities by issuing an initial gold standard problem task and constructs a state vector containing user capabilities, task history, and fluctuating states. Then, it uses a reinforcement learning agent to dynamically determine the timing and recipients of the gold standard problem task, optimizing the strategy with error reduction and gold standard problem cost as reward signals. Finally, it employs an iterative weighted aggregation method based on user capabilities and data errors to infer the task's truth value from the user's answers. This invention balances calibration accuracy and resource cost by dynamically optimizing the gold standard problem issuance strategy, addressing issues of excessive consumption or insufficient calibration. Simultaneously, it strengthens the aggregation algorithm's anti-interference capability to mitigate the influence of high noise and malicious users, effectively improving the robustness of truth inference.

[0050] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A resource-optimized dynamic crowd-sensing truth-finding method, characterized in that, The method includes the following steps: Step 1: Calculate the user's initial ability estimate based on the error between the user's actual result of completing the initial gold standard problem task and the known true value of the initial gold standard problem task. Step 2: Sort each user's initial capability estimate from highest to lowest, and assign a common perception task to each user according to the sorting result; wherein, the common perception task is completed by multiple users. Step 3: Iteratively calculate the perception data of all users participating in the same common perception task to obtain the final truth estimate of the current common perception task, and calculate the temporary capability index of each user based on the deviation between each user's perception data and the final truth estimate. Step 4: Construct a state vector based on the user's initial ability estimate and temporary ability index, and input the state vector into the reinforcement learning model to obtain the decision of the gold standard problem calibration task output by the reinforcement learning model, so as to assign the gold standard problem calibration task to the user based on the decision result; wherein, the reinforcement learning model includes a state vector, an action space and a reward function; Step 5: Generate a reward signal based on the user's capability estimate and the task release cost, and use the reward signal to update the parameters of the reinforcement learning model in order to optimize the generation of decisions based on the updated parameters.

2. The resource optimization dynamic crowd-aware truth discovery method according to claim 1, characterized in that, The expression for the user's initial capability estimate is: ; in, For users The initial capability estimate, The error between the user's measured value and the actual value from the calibration task. The preset upper limit of error, This is the lower limit threshold for the estimated user capability value.

3. The resource optimization dynamic crowd-aware truth discovery method according to claim 2, characterized in that, Step 3 specifically refers to: Step 3.1: Based on the users' initial ability estimates, obtain the initial weights for each user participating in the same general perception task, and normalize them. The expression is: ; in, For users The initial weights, The number of users participating in the same common perception task in this round. This is a user index variable that participates in the same general perception task in this round, with a value range of 1 to... This is used to refer to each user individually; Step 3.2: Calculate the weighted average based on the initial weights as the first aggregation result. The expression is: ; in, This is the initial aggregate estimate. For users The measured value; Step 3.3: Proceed In the next iteration, the absolute error between each user's measurement and the current aggregation result. for: ; in, For the first The aggregated estimate of the next iteration is weighted according to the magnitude of the error during the iteration process. The error weights are calculated as follows: ; in, For users In the The error weights obtained in the next iteration It is a constant; Step 3.4: Combine the error weights with the user capability estimates to obtain the comprehensive weight. The formula for calculating the comprehensive weight is: ; in, For users In the The comprehensive weights obtained in the nth iteration are normalized to obtain the nth iteration. Weight of +1 iteration The expression is: ; Step 3.5: Calculate the first... The result of +1 iteration yields a new aggregation result. The expression is: ; Step 3.6: Stop the iteration when the maximum relative change of the weights is less than the preset iteration convergence threshold. The expression is: ; in, This is the threshold for iterative convergence; Step 3.7: After the user completes the basic perception task, a temporary capability index is obtained based on the error between the user's measurement value and the final aggregated result. The expression is as follows: ; in, For users Temporary ability value, This represents the error between the user's measured value and the final aggregated result.

4. The resource optimization dynamic crowd-aware truth discovery method according to claim 1, characterized in that, The reinforcement learning model is specifically as follows: State vector S i = [ability, interval, variance, recent_mae], where ability is the user's current ability estimate, reflecting the quality level of the user's task completion; interval is the number of ordinary perception tasks the user has completed since the last completion of the gold standard problem task; variance is the standard deviation of the user's ability estimates over the last n times, reflecting the uncertainty or volatility of the ability estimates; and recent_mae is the average absolute error of the user in the last n ordinary perception tasks, reflecting the error between the user's perception data and the aggregated true value over the last n times. Action space A = {issue_GSQ, not_issue}, where issue_GSQ represents the gold standard issue to be published for the user, used to calibrate the user's capabilities; not_issue represents not publishing a gold standard issue for the user, and the user will continue to perform the ordinary perception tasks published by the platform. The reward function is reward = (base_reward + GSQ_total + stability_reward), where base_reward reflects whether the Gold Standard Question (GSQ) release decision improves the accuracy of user capability estimation; GSQ_total controls the effectiveness and frequency of GSQ use to ensure that GSQ releases are valuable; and stability_reward encourages the generation of stable and reliable capability estimates.

5. The resource optimization dynamic crowd-aware truth discovery method according to claim 4, characterized in that, The expression for base_reward is: ; in, The threshold is used to determine whether the adjustment range is within a predefined range. It is a natural constant. Estimate new capabilities and old capacity estimates The absolute difference is used to measure the magnitude of the adjustment; This is a positive reward coefficient used to adjust the reward intensity. This is a negative penalty coefficient used to adjust the intensity of the penalty. and The error attenuation coefficient is used to control the sensitivity of rewards and penalties as the adjustment range changes.

6. The resource optimization dynamic crowd-aware truth discovery method according to claim 4, characterized in that, The expression for GSQ_total is: ; in, The expression for the GSQ effect bonus is: ; in, This represents the effective reward coefficient for GSQ. This represents the GSQ invalidity penalty coefficient. and The intensity of rewards and penalties used to control the effectiveness of GSQ calibration; To effectively calibrate the error threshold, Invalid calibration error threshold, and The error threshold used to determine whether a GSQ is valid or invalid. The absolute difference between the new capability estimate and the old capability estimate; in, The expression for GSQ frequency reward is: ; in, To accumulate the number of GSQs published for users, , The threshold for GSQ frequency segmentation; To enable early use of reward coefficients, To avoid overuse of the penalty coefficient, and The intensity of rewards and penalties used to control the frequency of GSQ usage; This is the attenuation coefficient for excessive penalties.

7. The resource optimization dynamic crowd-aware truth discovery method according to claim 4, characterized in that, The expression for stability_reward is: ; in, The variance of the capability estimate is used to measure the stability of the user's capability. This serves as the base coefficient for stability rewards, used to control the upper limit of rewards for stability capabilities. This is the stability decay coefficient, used to control the variance of reward as a function of ability estimation. Increased decay rate.