Task Generation for ML Training Data via Worker Association
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current crowdsourcing methods for collecting big data suffer from significant variations in task processing results, which hinder the accuracy of machine learning due to workers with little knowledge about the training data being unable to provide correct label information, leading to increased time and potential rejection of tasks.
Innovation Solution
A task generation method that receives worker information, calculates degrees of association between analysis data and worker attributes, and extracts specific data for task processing, ensuring that workers with relevant knowledge are assigned tasks, thereby reducing variations in processing time and quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If tasks are distributed to a large number of workers through crowdsourcing, then data collection efficiency is improved, but variations in task processing quality increase
Solution Approach 1:
The patent applies local quality by matching specific workers to specific tasks based on their expertise and attributes. Instead of uniformly distributing tasks to all workers, the system calculates degrees of association between worker information (including expertise, location, language skills) and task requirements (such as data type, location information needed, language requirements), then assigns tasks to workers with the highest matching scores. This ensures that each task is handled by a worker with locally optimal qualifications for that specific task type.
Solution Approach 2:
The patent changes the parameter of worker-task matching from random or uniform distribution to association-based selection. The system calculates degrees of association using multiple parameters including worker expertise, location, language skills, and task requirements. By changing the assignment parameter from arbitrary to calculated association degrees, the system maintains high processing quality while distributing tasks across many workers.
2Adaptability or versatility
If workers with little knowledge about training data are assigned tasks, then task distribution coverage is improved, but task acceptance rates decrease
Solution Approach 1:
The patent applies preliminary action by pre-calculating and storing worker information including expertise, location, language skills, and other attributes before task assignment. The system maintains a worker database with these pre-processed characteristics, allowing rapid matching when tasks become available. This preliminary preparation enables the system to quickly identify suitable workers and assign tasks before workers reject them due to lack of relevant knowledge.
Solution Approach 2:
The system uses feedback from task processing results and worker performance to continuously improve task assignment accuracy. By monitoring which workers successfully complete tasks and producing high-quality results, the system refines its understanding of worker capabilities and adjusts future assignments accordingly, thereby maintaining high acceptance rates while expanding coverage.
3Ease of operation
If tasks are assigned without considering worker expertise, then assignment simplicity is improved, but processing time variations increase
Solution Approach 1:
The patent applies self-service by enabling the system to automatically perform task-worker matching without manual intervention. The automated calculation of association degrees between worker attributes and task requirements, followed by automatic task assignment, maintains operational simplicity while eliminating the need for manual expert matching. This automation resolves the contradiction by making the sophisticated matching process transparent and effortless for users.
Solution Approach 2:
The system changes the assignment parameter from simple random distribution to association-based selection, which reduces processing time variations by ensuring tasks are assigned to workers with appropriate expertise. The automated calculation of matching scores and subsequent assignment maintains ease of operation while significantly reducing the time workers need to spend understanding and processing tasks outside their knowledge domain.
Data Source
AI summary
A task generation method includes: receiving worker information from equipment of a worker over a network, the worker information including attribute information regarding a personal attribute of the worker; calculating degrees of association between each of pieces of analysis information resulting from analysis of pieces of data stored in a storage device connected to a computer and the worker information; extracting a piece of data to be subjected to task processing the worker is requested to perform from the pieces of data as specific data, based on the degrees of association; and generating a request task that is a task for making, to the equipment of the worker, a request for performing task processing for giving label information to the extracted specific data by using the equipment of the worker.


