Visitor Image-Response Training Data Generation Without Manual Annotation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AI training methods for inferring subjects in video require significant manual annotation effort from users, creating a heavy burden due to the large amount of data needed.
Innovation Solution
An information processing system that automatically generates training data using an intercommunication device to capture visitor images and responses, which are then processed by an inference server to determine subject reliability, reducing the need for manual annotation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual annotation work is performed to input audio type for each training data sample, then the quality and accuracy of training data can be ensured, but the user burden and time consumption increase significantly
Solution Approach 1:
The system automatically generates training data by utilizing existing intercommunication device logs and data without requiring manual user annotation. The data generation unit autonomously creates training datasets by combining audio data, image data, and response information from the intercommunication device, eliminating the need for users to manually input audio types for each sample.
Solution Approach 2:
The patent introduces an intercommunication device as an intermediary that automatically collects and stores structured data (audio recordings, images, response information) during normal operation. This intermediary device provides pre-processed, organized data that can be directly used for training without requiring additional manual annotation efforts.
2Measurement precision
If a large amount of training data is collected to improve AI model performance, then the accuracy of subject inference can be enhanced, but the requirement for manual annotation work increases proportionally
Solution Approach 1:
The system enables automatic generation of large-scale training datasets by leveraging data naturally accumulated from intercommunication device usage. The data generation unit continuously creates training samples from stored audio, image, and response data without requiring proportional increases in manual annotation resources, thus scaling data quantity without scaling annotation complexity.
Solution Approach 2:
The intercommunication device performs preliminary data collection and organization during normal operation, storing audio recordings, images, and response information in a structured format before they are needed for training. This preliminary action prepares the data in advance, eliminating the need for intensive post-collection annotation work when large datasets are required.
3Productivity
If automated data collection is implemented without manual annotation, then user burden is reduced and processing efficiency improves, but the quality and reliability of training data may deteriorate
Solution Approach 1:
The intercommunication device serves as a reliable intermediary that automatically collects and structures data during normal operation with built-in validation. The device stores audio data, image data, and response information in predetermined formats with inherent data quality controls, ensuring that automatedly generated training data maintains high reliability without requiring manual verification.
Solution Approach 2:
The system incorporates response information from the intercommunication device as feedback to verify and enhance training data quality. The response data provides contextual validation that the collected audio and image data correspond to actual communication events, ensuring the reliability of automatically generated training samples.
Data Source
AI summary
An information processing system includes an image capturing unit that captures an image of a visitor, an acquisition unit that acquires information about a response to a visit of the visitor, and a training data generation unit that generates training data for a learning model by using the image of the visitor captured by the image capturing unit and the information about the response acquired by the acquisition unit.


