Virtual Space Voice Data Routing With Audio-Visual Delivery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtual reality and augmented reality technologies lack effective methods for transmitting voice data between users in a virtual space, limiting interactive experiences.
Innovation Solution
A server system that extracts partial voice data from a user's utterance, determines a target user, and instructs their terminal to reproduce the data while displaying visual information based on additional voice data, enabling interactive communication in virtual environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If voice data is transmitted between users in virtual space, then interactive communication capability is improved, but system complexity increases
Solution Approach 1:
The patent segments voice data transmission into different types (first type and second type) with different processing methods. First type voice data is reproduced as audio, while second type voice data is converted to visual information. This segmentation allows the system to handle diverse communication needs without increasing overall complexity, as each segment has a dedicated processing path.
Solution Approach 2:
The server acts as an intermediary that receives voice data from users, determines the appropriate transmission type, and routes it to the target user's terminal. This intermediary approach centralizes the complex decision-making logic in the server, keeping client terminals relatively simple while enabling sophisticated interactive communication capabilities.
2Adaptability or versatility
If multiple types of voice data transmission are supported, then communication versatility is improved, but processing complexity increases
Solution Approach 1:
Voice data transmission is divided into two distinct types: first type that is reproduced as audio and second type that is converted to visual information. This segmentation enables the system to support multiple communication modes (audio-only, visual, or combined) while maintaining clear, separate processing paths for each type, thereby managing complexity through structured division.
Solution Approach 2:
The system changes the parameter of voice data representation by selecting between audio reproduction and visual information conversion based on the transmission type. This parameter change approach allows versatile communication modes without requiring complex processing for each mode, as the processing method is determined by a simple classification parameter.
3Loss of information
If visual information is displayed alongside voice data, then information completeness is improved, but device complexity increases
Solution Approach 1:
The patent segments information transmission into audio-based delivery and visual-based delivery. First type voice data is segmented for audio reproduction, while second type voice data is segmented for visual information conversion and display. This segmentation ensures complete information delivery through appropriate modalities without requiring a single device to handle all processing complexity.
Solution Approach 2:
The system adds a visual dimension to voice data transmission by converting second type voice data into visual information that is displayed on the terminal. This dimensionality change from purely audio to including visual representation provides more complete information while keeping the processing approach simple: either audio reproduction or visual conversion based on transmission type.
Data Source
AI summary
An example server for constructing a virtual space includes a memory configured to store computer-executable instructions and a processor configured to execute the instructions by accessing the memory. The instructions, when executed, cause the processor to extract first partial voice data corresponding to a target utterance from voice data of a first user received from a terminal of the first user among users in the virtual space; instruct a target terminal of the target user to reproduce the first partial voice data; and, based on transmission of second partial voice data of a second user to the target user being requested while the target terminal reproduces the first partial voice data, instruct the target terminal to display visual information generated based on the second partial voice data.


