Voice-Image Fusion for Unauthorized Access Prevention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Artificial intelligence speakers face issues with incorrectly recognizing sounds from radios or TVs as voice commands and security vulnerabilities due to unauthorized voice commands from unverified users.
Innovation Solution
An electronic device that combines voice and image information to verify the authenticity of voice commands by capturing images of the subject and using face recognition algorithms to determine if the voice is genuine and from a registered user, thereby controlling device operations accordingly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If voice recognition is used for device control, then ease of operation is improved, but security is worsened due to unauthorized voice commands
Solution Approach 1:
The patent combines voice recognition with face recognition technology to create a dual verification system. The voice sensor detects voice commands while the camera captures facial images, and the controller integrates both modalities to verify user identity. This merging of sensing modalities maintains voice control convenience while significantly improving security by ensuring the voice command comes from an authorized user.
Solution Approach 2:
The controller acts as an intermediary that mediates between voice detection and execution of control commands. Before executing a voice command, the controller first verifies the speaker's identity through face recognition using captured images. This intermediary verification step prevents unauthorized users from executing commands even if they successfully trigger voice recognition.
2Device complexity
If voice recognition alone is used, then device complexity is reduced, but measurement precision is worsened due to inability to verify user identity
Solution Approach 1:
The patent merges voice recognition with face recognition to improve measurement precision of user identity verification. The voice sensor captures acoustic signals while the camera captures visual facial images, providing complementary information that together enable accurate verification of user identity, overcoming the limitations of voice recognition alone.
Solution Approach 2:
The patent adds another dimension of verification by incorporating facial image recognition alongside voice recognition. Instead of relying on a single modalities, the system verifies identity through both acoustic (voice) and visual (face) dimensions, significantly improving the precision of user identification while maintaining manageable system complexity through integrated processing.
3Reliability
If face recognition is added to voice recognition, then security is improved, but device complexity increases
Solution Approach 1:
The controller is designed to perform multiple functions: processing voice commands, capturing facial images, executing face recognition, and integrating both verification modalities. By making the controller universal and multi-functional, the patent manages device complexity through software-based integration rather than requiring separate dedicated hardware systems for each function.
Solution Approach 2:
The patent merges the face recognition and voice recognition systems into a unified verification framework within the controller. The camera and voice sensor work together under coordinated control, with the controller integrating their outputs to make verification decisions. This merging approach improves security through multi-modal verification while managing complexity through centralized processing.
Data Source
AI summary
The present invention includes: a voice sensor for detecting voice information; a camera for capturing an image of a subject related to the voice information; and a control unit for controlling the camera such that the image of the subject related to the voice information is captured when the voice sensor detects the voice information, and determining, by using the captured image of the subject and the voice information, whether the subject related to the voice information is a counterfeit face, thereby determining whether to execute a control command corresponding to the voice information.


