Audio-Visual Reality Chatbot for Real-Time Environmental Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing natural language chatbots like ChatGPT lack the ability to provide real-time, scenario-specific responses due to their reliance on pre-learned data, failing to adapt to the user's current environment and needs.

Innovation Solution

A system and method that utilizes an audio-visual reality interface to capture environment images and sounds, integrating a chatbot that runs a natural language model to generate dialogue content based on user location, preferences, and real-time environmental data, enabling personalized and contextually relevant interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a natural language chatbot is used to provide general discussions, then the chatbot can respond to users in natural language, but the responses are standard answers obtained through learning and cannot be adapted in real-time to provide answers relevant to the current status of the user

Engineering Contradiction:
Improveadaptability to real-time scenariosVSAvoidlack of real-time environmental context
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent introduces an intermediary system that includes a reality image acquisition module, object recognition module, and audio recognition module. These modules act as mediators between the user's physical environment and the chatbot, extracting environmental information and transmitting it to the chatbot processing system. This intermediary layer enables the chatbot to access real-time environmental context without directly modifying the core chatbot architecture.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent enhances the chatbot's universality by integrating multiple functional modules: reality image acquisition, object recognition, audio recognition, and location-based services. The chatbot system can now perform multiple functions including general conversation, environmental analysis, object identification, and context-aware response generation, making it adaptable to diverse real-time scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If the application is not integrated with the actual environment, then the chatbot can maintain simple operation, but effective responses cannot be provided based on a current scenario of the user

Engineering Contradiction:
Improvescenario-specific response capabilityVSAvoidintegration with actual environment
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the system into distinct functional modules: a chatbot interface module for natural language processing, a reality image acquisition module for capturing environmental data, an object recognition module for identifying objects, an audio recognition module for capturing sounds, and a location-based service module for geospatial information. Each module operates independently and communicates through standardized interfaces, reducing overall system complexity while enabling comprehensive environmental integration.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a nested architecture where smaller functional modules are contained within the larger chatbot system. The object recognition module, audio recognition module, and location-based service module are nested within the chatbot application framework, allowing each component to be developed, tested, and updated independently while contributing to the overall scenario-specific response capability.

Inventive Principle:
Principle #7Nested doll (Nesting)

3Adaptability or versatility

If the chatbot uses pre-learned data for responses, then the chatbot can provide consistent answers, but the responses cannot be personalized to user preferences and current location

Engineering Contradiction:
Improvepersonalization to user and locationVSAvoidtime for real-time data processing
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-processing and storing user profile information, preferences, and location data in databases before the actual conversation occurs. The system pre-loads relevant contextual information and maintains ready-to-access data structures that can be quickly retrieved during real-time interaction, reducing the time needed for personalization while enabling tailored responses based on user preferences and current location.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250233835A1Method and system for triggering an intelligent dialogue through an audio-visual reality
Publication Date: 2025.07.17 PLAYSEE INC
  • US20250233835A1 patent drawing
  • US20250233835A1 patent drawing
  • US20250233835A1 patent drawing

AI summary

A method and a system for triggering an intelligent dialogue through an audio-visual reality. In the method, a reality image interface is initiated in a user device, a camera is activated to obtain an environment image, and a microphone is activated to obtain an environment audio. Then, a cloud server receives location information and a reality image request from the user device, obtains an environment object by identifying the environment image and an environment sound by identifying the environment audio. Afterwards, an intelligent dialogue link point displayed on the reality image interface is triggered to activate an intelligent dialogue program. An intelligent dialogue interface is initiated in the intelligent dialogue program and a chatbot is introduced in the intelligent dialogue program. The chatbot runs a natural language model to generate a dialogue content based on the location information, the environment object, and the environment sound.