Gameplay Telemetry Captioning via Neural Network

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional visual description techniques for providing captions for video games are labor-intensive, time-consuming, and costly, mainly relying on manual creation of descriptive transcripts and human voice actors, which limits their application to pre-recorded content.

Innovation Solution

A data processing apparatus and method that utilize an artificial neural network (ANN) to generate caption data based on gameplay telemetry data, eliminating the need for video image processing and enabling real-time or recorded video game session captioning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual creation of descriptive transcripts and human voice actors are used, then visual description accuracy is improved, but labor intensity and cost increase

Engineering Contradiction:
Improvevisual description accuracyVSAvoidlabor efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces the mechanical system of manual transcription and human voice acting with an automated computer-based system that captures gameplay telemetry data and generates captions programmatically, eliminating the need for human labor while maintaining description accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by automatically generating visual descriptions through telemetry data processing without requiring human intervention for transcription or voice synthesis, allowing the system to serve itself in creating accessible content

Inventive Principle:
Principle #25Self-service

2Measurement precision

If manual transcription methods are used, then description quality is improved, but time consumption increases

Engineering Contradiction:
Improvedescription qualityVSAvoidcaption generation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by capturing and storing gameplay telemetry data during the game session itself, so that when captions are needed, the data is already available and processed, eliminating the need for time-consuming post-processing transcription

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent substitutes the time-intensive manual transcription process with automated computational processing of telemetry data, generating captions instantly without human intervention

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If video image processing is used for caption generation, then visual accuracy is improved, but computational load and latency increase

Engineering Contradiction:
Improvevisual accuracyVSAvoidcomputational load
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the necessary gameplay information through telemetry data collection during game execution, avoiding the need to process entire video images while still capturing all relevant visual events and actions for accurate captioning

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system segments the visual information extraction process by using dedicated telemetry data collection for specific game events rather than processing complete video frames, reducing computational overhead while maintaining accuracy

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If conventional visual description techniques are used, then description accuracy is improved, but adaptability to live content decreases

Engineering Contradiction:
Improvedescription accuracyVSAvoidapplicability to live content
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system implements dynamics by enabling real-time caption generation during live gameplay sessions through continuous telemetry data processing, allowing the system to adapt to both recorded and live content dynamically rather than being limited to pre-recorded material

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250177858A1Apparatus, systems and methods for visual description
Publication Date: 2025.06.05 SONY INTERACTIVE ENTERTAINMENT LLC
  • US20250177858A1 patent drawing
  • US20250177858A1 patent drawing
  • US20250177858A1 patent drawing

AI summary

A data processing apparatus comprises a captioning model to receive gameplay telemetry data indicative of one or more in-game properties for a session of a video game, the captioning model comprising an artificial neural network (ANN) trained to output caption data comprising one or more captions in dependence upon a learned mapping between gameplay telemetry data and caption data, one or more of the captions comprising one or more words for providing a visual description for the session of the video game, and output circuitry to output one or more of the captions.