Scene Estimation Using Position-Aware Multi-Signal Feature Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As the number of acoustic and video signals used for scene estimation increases, the curse of dimensionality occurs, leading to reduced accuracy despite the increased amount of information, making it difficult to accurately estimate scenes.

Innovation Solution

The scene estimation device integrates acoustic and video feature amounts using conditional encoders that consider the signal acquisition positions, reducing feature dimensionality through neural networks like multi-layer CNNs and ResNet, followed by scene selection using integrated feature amounts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If the number of acoustic signals and video signals used for scene estimation increases, then the amount of information for scene estimation increases, but the dimensionality of data increases causing the curse of dimensionality and reducing estimation accuracy

Engineering Contradiction:
Improveamount of information for scene estimationVSAvoidscene estimation accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent extracts only the necessary features from acoustic and video signals using dedicated encoders (acoustic encoder, video encoder, and conditional encoder). Instead of using all raw signal data, the system extracts essential feature amounts that capture scene information while reducing dimensionality, thereby avoiding the curse of dimensionality while preserving estimation accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the scene estimation process into multiple independent encoding stages: acoustic signal encoding, video signal encoding, and conditional encoding based on acquisition positions. Each encoder processes specific input signals separately and generates intermediate feature amounts that are then integrated, allowing the system to handle multiple signals without overwhelming dimensionality.

Inventive Principle:
Principle #1Segmentation

2Loss of information

If the number of input acoustic signals and video signals increases, then blind spots are reduced and information completeness improves, but data dimensionality increases leading to higher computational complexity

Engineering Contradiction:
Improveinformation completeness for scene estimationVSAvoidcomputational complexity of scene estimation processing
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts essential feature amounts from multiple acoustic and video signals using efficient encoding processes. The acoustic encoder, video encoder, and conditional encoder work together to extract only the most relevant information from each signal, reducing the overall data dimensionality while maintaining information completeness for accurate scene estimation.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent merges the encoding of acoustic signals, video signals, and acquisition position information through a conditional encoder. This integration allows the system to process multiple input signals in a unified manner, combining their feature amounts to represent the overall scene while avoiding the computational complexity of processing each signal separately at full dimensionality.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If more acoustic signals and video signals are used for scene estimation, then coverage of the scene improves, but the dimensionality of handled data increases causing accuracy degradation

Engineering Contradiction:
Improvescene estimation reliabilityVSAvoidscene estimation accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent extracts essential scene-relevant features from multiple acoustic and video signals while discarding redundant information. The conditional encoder specifically extracts features related to acquisition positions, and the integration of these extracted features maintains scene estimation reliability without suffering from the dimensionality curse that would otherwise degrade accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12609133B2Scene selection method, scene selection apparatus and program
Publication Date: 2026.04.21 NT T INC
  • US12609133B2 patent drawing
  • US12609133B2 patent drawing
  • US12609133B2 patent drawing

AI summary

Provided is a technique for accurately estimating a scene even when the number of input signals increases. A scene estimation method includes: when S is the number of scenes and M is the number of input acoustic signals, an acoustic signal encoding step of generating, by a scene estimation device, an integrated acoustic feature amount from an m-th input acoustic signal (m=1, . . . , M) and a position where the m-th input acoustic signal is acquired (hereinafter referred to as an m-th input acoustic signal acquisition position) (m=1, . . . , M); and a scene selection step of selecting, by the scene estimation device, a scene from which M input acoustic signals are acquired from among S scenes, using the integrated acoustic feature amount.