Digital Camera Voice Recognition Face Linking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Digital cameras cannot identify or display the faces of individuals whose voices are captured in images or videos, especially in moving images, where voices of the photographer and surrounding persons can be heard but their identities remain unknown.
Innovation Solution
A digital photographing apparatus equipped with a digital signal processor (DSP) that recognizes voices and matches them to corresponding faces, storing this information in an EXIF format, allowing the DSP to display the face on the screen, blink it in sync with the voice, and determine the sound source direction for accurate positioning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If voice recognition is added to identify voice owners, then the ability to identify persons whose voices are captured is improved, but the device complexity increases due to additional voice recognizing unit and processing requirements
Solution Approach 1:
The patent combines voice recognition functionality with the existing face recognition and image processing system. The voice recognizing unit is integrated into the digital signal processor that already handles face detection and image data processing, allowing voice owner identification to be merged with existing functions rather than adding completely separate systems
Solution Approach 2:
The digital signal processor is designed to perform multiple functions including face detection, face recognition, voice recognition, and sound source direction determination. By making the processor universal and multi-functional, the patent avoids the need for separate dedicated hardware for each function, thereby managing device complexity while achieving voice owner identification
2Loss of information
If the face corresponding to recognized voice is displayed on the screen, then the visibility of voice owner is improved, but the information processing load increases requiring additional storing unit and comparison operations
Solution Approach 1:
The patent performs preliminary face recognition and voice recognition operations during image capture, storing the recognized voice and corresponding face information in advance in the storing unit. This preliminary processing allows for quick retrieval and comparison during playback without requiring heavy real-time processing, thus managing information processing load effectively
Solution Approach 2:
The patent creates and stores copies of face images and voice data in the storing unit for later retrieval and comparison. By pre-storing these data copies, the system avoids the need for complex real-time analysis during playback, reducing the information processing load while enabling voice owner visibility through face display
3Measurement precision
If sound source direction analysis is implemented to position the face accurately, then the positioning precision is improved, but the measurement and detection difficulty increases due to additional direction determining unit
Solution Approach 1:
The patent utilizes the existing microphone array and signal processing capabilities of the digital photographing apparatus to perform sound source direction determination. The system uses its own built-in hardware and processing resources rather than requiring external or additional specialized equipment, thereby achieving improved positioning precision without proportionally increasing detection difficulty
Data Source
AI summary
A digital photographing apparatus and a control method thereof are disclosed. The digital photographing apparatus receives a voice, recognizes the received voice, and stores the recognized voice to correspond to a face. In this way, by recognizing the voice while displaying an image, an owner of the voice can be identified and thus a face of a person which is not included in an image being captured can be displayed on a screen.


