Drone Swarm Multi-View Capture for First-Person Scene Reconstruction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional multi-view cinematography faces challenges in capturing immersive first-person views, especially in outdoor scenes or large spaces, due to restrictions on camera placement and control, requiring precise alignment and complex post-processing, and lacks automated systems for tracking targets without physical contact or prior knowledge of movement.

Innovation Solution

A system comprising two swarms of drones, where the first swarm captures images of a target's face to determine head pose and gaze, and transmits this information to the second swarm to capture environmental images, allowing for real-time tracking and generation of first-person views without physical contact or prior knowledge of the target's movement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If cameras are mounted on the actor's head to capture first-person view, then first-person view can be captured, but the actor's movement is restricted and the performance naturalness is affected

Engineering Contradiction:
Improveactor movement freedomVSAvoidcamera mounting complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical camera mounting system on the actor's head with an automated drone swarm system that uses vision-based tracking and computer algorithms to capture first-person views from a distance, eliminating the need for physical contact with the actor

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces drones as an intermediary between the camera system and the actor, allowing cameras to capture first-person views without being mounted on the actor's head, thus avoiding movement restrictions and performance interference

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If multiple fixed cameras are used for multi-view cinematography, then multi-view images can be captured, but the system cannot track targets in large spaces or outdoor scenes

Engineering Contradiction:
Improvespatial coverageVSAvoidcamera system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent transforms the static fixed camera system into a dynamic drone swarm system that can move freely through large spaces and outdoor environments, enabling multi-view cinematography in previously inaccessible locations

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent divides the camera system into multiple independent drone units, each capable of autonomous flight and image capture, allowing the system to cover large spatial areas that would be impossible for a single fixed camera system

Inventive Principle:
Principle #1Segmentation

3Object-affected harmful factors

If drones are used to track targets autonomously, then physical contact with the target is avoided, but computationally intensive scene analysis is required during filming

Engineering Contradiction:
Improvephysical contact with targetVSAvoidcomputational energy consumption
Core Design Contradiction:
Object-affected harmful factorsVSUse of energy by moving object

Solution Approach 1:

The patent performs target detection and tracking setup in advance using the first plurality of drones positioned in front of the target, determining head pose and gaze before the second plurality of drones captures the final images, reducing real-time computational demands

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4136618B1System of multi-view image capturing using drone swarms
Publication Date: 2025.03.26 SONY GROUP CORP
  • EP4136618B1 patent drawingFigure 1
  • EP4136618B1 patent drawingFigure 2
  • EP4136618B1 patent drawingFigure 3

AI summary

A system of multi-view imaging of an environment through which a target moves includes pluralities of drones, each drone having a drone camera. A first plurality of drones moves to track target movement, capturing a corresponding first plurality of images of the target's face, making real time determinations of the target's head pose and gaze from the captured images and transmitting the determinations to a second plurality of drones. The second plurality of drones moves to track target movement of the target, with drone camera poses determined at least in part by the head pose and gaze determinations received from the first plurality of drones, in order to capture a second plurality of images of portions of the environment in front of the target. Post-processing of the second plurality of images allows generation of a first-person view representative of a view of the environment seen by the target.