Video Conversational Interface Transition Masking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video-based conversational interfaces using pre-recorded video segments face challenges in providing a flexible and engaging user experience due to noticeable transitions between video clips, which can result in a visually jarring experience.
Innovation Solution
A method that transitions between pre-recorded video segments by enlarging the user's self-video image to mask the change, using a peripheral display region for the previous video segment and a main display region for the new segment, with incremental opacity and area changes to minimize visual disruption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If pre-recorded video segments are used for conversational responses, then user engagement and realism are improved, but frequent transitions between video clips create visually jarring experiences
Solution Approach 1:
The user's self-video image serves as an intermediary element that masks the transition between different pre-recorded video segments. By positioning the self-video in the peripheral display region and gradually changing its opacity, the system mediates the visual switch between video clips, making the transition imperceptible to users.
Solution Approach 2:
The system changes the opacity (visual property) of the self-video image gradually during transitions. By adjusting the transparency of the self-video overlay, the system creates a smooth visual transition that masks the underlying video segment changes, reducing visual disruption.
2Ease of operation
If video segments are transitioned at predetermined points, then conversation flow is maintained, but transitions occur at noticeable moments creating visual disruption
Solution Approach 1:
The self-video image acts as a mediator that obscures the main display region during transition periods. This allows the system to switch video segments at any point without users noticing, as the self-video overlay blocks the visual change while maintaining conversation continuity.
Solution Approach 2:
The system prepares transition masks (using the self-video image) in advance and gradually introduces them before video segment changes occur. This preliminary action of fading in the self-video overlay ensures that when the actual video transition happens, users are already visually adapted and won't notice the switch.
3Object-affected harmful factors
If the self-video image is continuously displayed in the main region, then transitions are masked effectively, but the conversational interface becomes less engaging
Solution Approach 1:
The system moves the self-video image from the main display region to the peripheral display region during non-transition periods. This spatial repositioning in another dimension (from central to peripheral) allows the main video content to remain prominent and engaging, while the self-video is still available to mask transitions when needed.
Solution Approach 2:
The self-video image dynamically changes position and opacity based on system state. It appears in the peripheral region during normal conversation, then moves to mask transitions, then returns to its original position. This dynamic behavior ensures both engagement during conversation and effective masking during transitions.
Data Source
AI summary
In an answer view, a first video segment is selected based on a first natural language input and displayed in a main display region, and a self-video image of a user is displayed in a peripheral display region having a smaller area than the main display region. To transition from the answer view to a question view, the self-video image is enlarged to replace the first video segment in the main display region. A second natural language input is received. To transition from the question view to the answer view, the self-video image is reduced to occupy the peripheral display region and the self-video image is replaced in the main display region with a second video segment selected based on the second natural language input. The video segments are pre-recorded video response segments spoken by the same person. Enlarging the self-video image masks the transition between the video segments.


