Video Game Voice Selection Using Video Content Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video games face challenges with voice-based communication, particularly in multiplayer scenarios, where simultaneous voice inputs from multiple users can result in unintelligible mixtures of sounds, leading to user overburdening and hindered communication.
Innovation Solution
A data processing apparatus and method that selects a subset of speech signals relevant to the video game content by analyzing video images and comparing them with speech inputs, using computer vision and natural language processing to filter out irrelevant or disruptive communications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If voice-based communication is provided for multiple users simultaneously, then user interaction and coordination are enhanced, but communication intelligibility deteriorates due to overlapping sounds
Solution Approach 1:
The system extracts and selects only the most relevant speech signals from among multiple simultaneous user inputs. By using video analysis to identify relevant actions and comparing them with speech content, the system extracts only those speech signals that correspond to relevant in-game events, filtering out irrelevant or disruptive communications while preserving useful interactions.
Solution Approach 2:
The system applies different quality levels to different speech signals based on their relevance. Speech signals associated with relevant video content are maintained with full quality, while irrelevant signals are filtered out. This local differentiation ensures that only the most important communications are delivered to users, improving overall intelligibility without sacrificing useful interactions.
2Loss of information
If all user voice communications are output simultaneously, then complete information is provided, but user overburdening increases due to excessive audio inputs
Solution Approach 1:
The system extracts only the necessary speech signals from the complete set of user inputs. By analyzing the relationship between video content and speech, it identifies and extracts only those communications that are relevant to current in-game actions, thereby reducing the total audio load on users while maintaining information completeness for relevant events.
Solution Approach 2:
Instead of providing all possible speech signals (excessive action), the system provides only the partial set that is actually relevant to current video content. This partial action approach prevents user overburdening by filtering out redundant or irrelevant communications while still ensuring that all important information is conveyed.
3Adaptability or versatility
If voice-based communication is enabled for large numbers of users, then interaction opportunities increase, but communication becomes infeasible due to simultaneous speaking
Solution Approach 1:
The system extracts and delivers only the most relevant speech signals to each user based on their current in-game actions and the corresponding video content. This extraction process maintains communication feasibility even with large numbers of users by ensuring that each user receives only the speech signals that are relevant to their current activities, preventing the chaos of simultaneous speaking.
Solution Approach 2:
The system uses feedback from video analysis to dynamically adjust which speech signals are delivered to which users. By continuously monitoring video content and comparing it with speech inputs, the system creates a feedback loop that adapts the communication delivery to maintain feasibility and relevance, even as the number of interacting users varies.
Data Source
AI summary
A data processing apparatus comprises receiving circuitry to receive video images and associated speech signals for a video game, the speech signals indicative of speech input for a plurality of users associated with the video game, analysis circuitry to analyse at least some of the video images and generate video description data indicative of one or more properties for the video images, and selection circuitry to select a subset of the speech signals to be output for the video game, wherein the selection circuitry is configured to select a respective speech signal responsive to whether a comparison for the respective speech signal and the video description data satisfies a selection condition.


