Multi-Camera 3D Teleconference Calibration with Foundation Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing 3D teleconferencing systems face challenges in generating high-quality 3D models due to variations in subject attributes, camera properties, and environmental factors, leading to inaccuracies in depth perception and representation.

Innovation Solution

A machine learning model is trained to infer optimal camera settings based on subject attributes, camera properties, and environmental conditions, automatically adjusting camera pose and other settings to improve 3D model fidelity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual adjustments are made to camera settings during calibration, then 3D model quality improves, but time consumption and operational complexity increase

Engineering Contradiction:
Improve3D model qualityVSAvoidcalibration time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by training a machine learning model on historical calibration data and subject attributes before actual 3D scanning. The pre-trained model automatically infers optimal camera settings based on input subject characteristics, eliminating the need for manual calibration adjustments during each scanning session while maintaining high 3D model quality.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If multiple camera settings are adjusted to account for subject attributes, then measurement precision improves, but device complexity increases

Engineering Contradiction:
Improvedepth perception accuracyVSAvoidcamera configuration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system implements self-service by using a machine learning model that automatically analyzes subject attributes (such as skin tone, size, and shape characteristics) and independently determines the optimal camera settings. The model receives input about the subject and autonomously configures camera pose, focus depth, and white balance without requiring manual intervention or complex configuration procedures.

Inventive Principle:
Principle #25Self-service

3Manufacturing precision

If camera settings are optimized for specific subject attributes, then 3D model fidelity improves, but adaptability to different subjects decreases

Engineering Contradiction:
Improve3D model fidelityVSAvoidsubject compatibility
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The system applies parameter changes by using a machine learning model that dynamically adjusts camera settings based on varying subject attributes. The model is trained on diverse data representing different skin tones, subject sizes, and environmental conditions, enabling it to generalize across various subjects. When a new subject is scanned, the model analyzes their specific attributes and automatically configures appropriate camera parameters, maintaining high 3D model fidelity across different subject types.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250259380A1System for optimizing sensor settings in a multi-camera environment based on foundation models and historical data
Publication Date: 2025.08.14 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250259380A1 patent drawing
  • US20250259380A1 patent drawing
  • US20250259380A1 patent drawing

AI summary

3D teleconferences use an array of cameras to generate a 3D model of a subject. During a calibration and registration process the pose of each camera may be adjusted. Similarly, camera settings such as focus depth and white balance may be modified. These changes are made to improve the quality of the 3D model generated from image data captured by the cameras. Many factors affect the quality of images captured by the cameras. For example, depth sensors may be affected by the skin tone of the subject. In some configurations, a machine learning model (ML model) is trained on adjustments to properties that affect 3D model quality. The resulting ML model may then be used to infer camera adjustments for a given set of subject attributes, camera properties, and/or environment properties.