Audio-Visual Reality Chatbot for Real-Time Contextual Dialogue

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing natural language chatbots like ChatGPT lack real-time adaptability to provide user-relevant responses due to standard answers based on learned data, failing to integrate with the user's current scenario effectively.

Innovation Solution

A system and method that utilizes an audio-visual reality interface to capture environment images and sounds, leveraging a cloud server with chatbots trained on machine learning and NLP to generate dialogue content that matches user preferences and real-time scenarios by integrating location-based data and user activities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If standard answers from learned data are used, then the chatbot can respond in natural language, but the responses cannot be adapted in real-time to provide answers relevant to the current status of the user

Engineering Contradiction:
Improvereal-time adaptabilityVSAvoiduser context information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system pre-integrates multiple information gathering mechanisms (camera, microphone, location services, sensor inputs) that continuously collect environmental data before user queries occur. This preliminary action ensures context information is already available when the user asks a question, enabling real-time adaptation without delay.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary layer between the user and the chatbot that processes environmental context information. This intermediary analyzes camera images, audio inputs, location data, and sensor readings to extract relevant context, then feeds this processed information to the chatbot system, enabling it to adapt responses based on current user status.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the application is not integrated with the actual environment, then the system structure remains simple, but effective responses cannot be provided based on a current scenario of the user

Engineering Contradiction:
Improveenvironmental integrationVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system employs a multi-functional architecture where a single chatbot interface can handle various types of inputs (text, voice, images, location data, sensor information) and provide context-aware responses across different scenarios. This universal design allows the system to integrate multiple environmental data sources without proportionally increasing complexity, as the same core processing mechanisms handle diverse input types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements a nested structure where environmental context information (camera data, audio, location, sensor inputs) is layered within the chatbot processing system. The chatbot core contains the fundamental NLP capabilities, while environmental context layers are nested around it, each contributing additional information without disrupting the core functionality. This nesting allows gradual integration of complexity.

Inventive Principle:
Principle #7Nested doll (Nesting)

3Measurement precision

If location-based data and environmental information are integrated, then dialogue content can match user scenario, but the system requires multiple sensors and data processing components

Engineering Contradiction:
Improvescenario matching accuracyVSAvoidsensor and processing components
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system merges multiple data processing functions into a unified context analysis module that simultaneously processes camera images, audio inputs, location data, and sensor readings. Instead of having separate processing pipelines for each sensor type, the patent combines them into an integrated system that analyzes all environmental inputs together, extracting context information through a single coordinated process that reduces overall system complexity.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP4589449A1Method and system for triggering an intelligent dialogue through an audio-visual reality
Publication Date: 2025.07.23 PLAYSEE INC
  • EP4589449A1 patent drawingFigure 1
  • EP4589449A1 patent drawingFigure 2
  • EP4589449A1 patent drawingFigure 3

AI summary

A method and a system for triggering an intelligent dialogue through an audio-visual reality. In the method, a reality image interface is initiated in a user device (150), a camera is activated to obtain an environment image, and a microphone (703) is activated to obtain an environment audio. Then, a cloud server (100) receives location information and a reality image request from the user device (150), obtains an environment object by identifying the environment image and an environment sound by identifying the environment audio. Afterwards, an intelligent dialogue link point displayed on the reality image interface is triggered to activate an intelligent dialogue program. An intelligent dialogue interface is initiated in the intelligent dialogue program and a chatbot is introduced in the intelligent dialogue program. The chatbot runs a natural language model to generate a dialogue content (801) based on the location information, the environment object, and the environment sound.