Automated Audio Prompt Selection via Object Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing home security systems require occupants to directly interact with visitors through audio/video recording and communication devices, limiting convenience when the occupant is unavailable or engaged, such as during a movie or business meeting.

Innovation Solution

Network-connected security devices, like A/V recording and communication devices, utilize automated audio prompts played through speakers based on visitor identification, using object recognition to provide instructions without direct occupant interaction, and allow users to select or customize prompts for playback through client devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If automated audio prompts are implemented, then convenience is improved by allowing automated instruction delivery without occupant interaction, but device complexity increases due to the need for object recognition and automated prompt selection systems

Engineering Contradiction:
ImproveconvenienceVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces an automated audio prompt system that acts as an intermediary between the occupant and the visitor. The system includes a processor that receives image data, identifies objects/visitors, selects appropriate audio prompts from a database, and plays them through speakers. This intermediary automation resolves the contradiction by handling the instruction-delivery function without requiring direct occupant involvement, thereby improving convenience while managing complexity through modular system design.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The automated audio prompt system enables the doorbell device to serve itself by automatically identifying visitors through image recognition, selecting appropriate prompts from stored options, and delivering instructions without human intervention. The processor autonomously manages the entire workflow from receiving image data to playing audio prompts, allowing the system to perform its communication function independently and improving ease of operation.

Inventive Principle:
Principle #25Self-service

2Reliability

If object recognition is used to identify visitors, then communication effectiveness is improved by providing targeted instructions, but measurement precision requirements increase for accurately identifying and categorizing visitors

Engineering Contradiction:
Improvecommunication effectivenessVSAvoididentification accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent applies partial action by implementing a tiered identification approach where the system processes image data to identify visitors at varying levels of detail. The processor can recognize basic object categories (person, package, animal) and select appropriate prompts based on these partial identifications, rather than requiring complete and precise identification of every visitor attribute. This allows the system to achieve sufficient communication effectiveness without demanding excessive measurement precision.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The audio prompt database is segmented into multiple categories corresponding to different visitor types (e.g., delivery persons, family members, strangers, packages, animals). The processor segments the identification task by matching detected objects to these predefined categories, selecting prompts based on the best match rather than requiring perfect identification accuracy. This segmentation approach improves reliability by ensuring appropriate prompts are delivered even when identification is not perfectly precise.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11736665B1Custom and automated audio prompts for devices
Publication Date: 2023.08.22 AMAZON TECH INC
  • US11736665B1 patent drawing
  • US11736665B1 patent drawing
  • US11736665B1 patent drawing

AI summary

A network-connected security device is communicatively coupled to an audio/video (A/V) recording and communication device having a camera and a speaker. A method receives video data captured by the camera, and performs an object recognition algorithm upon the received video data to identify an object therein. The method performs a table lookup using the identified object, into a data structure that associates objects with at least one description of a predefined voice message. The method selects a description of a predefined voice message associated with the identified object, and transmits the selected description's predefined voice message to the A/V recording and communication device for output through the speaker.