Camera Footage Abstraction for Privacy and Smaller Scene Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing camera footage management systems fail to adequately protect human privacy and efficiently reduce data size, particularly due to incorrect determination of image processing ranges and increased data size from added text information.

Innovation Solution

A system that generates recognition and linguistic information from camera footage, renders abstracted images of humans and spaces, and adds caption information, reducing data size while maintaining privacy through abstracted representations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If text information is added to output video to notify abnormal scenes, then notification effectiveness is improved, but data size of video increases

Engineering Contradiction:
Improvenotification effectivenessVSAvoiddata size
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent extracts text information from video frames and stores it separately as metadata rather than embedding it directly into the video data. This allows the text information to be preserved for notification purposes while keeping the video data size minimal.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system uses an intermediary metadata structure that links video frames with their corresponding text information and abnormal scene indicators. This mediator allows efficient access to notification information without requiring storage of large video datasets.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If privacy protection processing is applied to surveillance footage, then human privacy is protected, but range determination accuracy deteriorates

Engineering Contradiction:
Improveprivacy protectionVSAvoidrange determination accuracy
Core Design Contradiction:
Object-affected harmful factorsVSManufacturing precision

Solution Approach 1:

The patent segments the video processing into two distinct stages: first, analysis of the original video to determine abnormal scenes and generate text information; second, application of privacy protection processing to the output video. This segmentation allows accurate range determination in the analysis stage while ensuring privacy protection in the output stage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary analysis of the video content to identify abnormal scenes and generate descriptive text information before applying privacy protection processing. This preliminary action ensures that accurate scene understanding is achieved before any privacy-obscuring transformations are applied.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If score table with fine granularity is used for importance calculation, then importance accuracy is improved, but range determination correctness worsens

Engineering Contradiction:
Improveimportance calculation accuracyVSAvoidrange determination correctness
Core Design Contradiction:
Measurement precisionVSManufacturing precision

Solution Approach 1:

Instead of using detailed score tables to determine processing ranges directly, the patent inverts the approach by using linguistic processing to generate text descriptions of scenes, and then using these text descriptions to determine which ranges require privacy protection. This inversion avoids the pitfalls of overly granular scoring while maintaining accuracy.

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS20250356568A1Camera footage management system
Publication Date: 2025.11.20 TOYOTA JIDOSHA KK
  • US20250356568A1 patent drawing
  • US20250356568A1 patent drawing
  • US20250356568A1 patent drawing

AI summary

A processing circuitry generates recognition information on an object shown in camera footage by object recognition processing on the camera footage. The processing circuitry also generates linguistic information on a scene shown in the camera footage by linguistic processing on the camera footage and generates scene information in which the recognition information on a human and the linguistic information on the scene are associated with each other and stores the scene information in the memory device. The processing circuitry further performs reproduction processing of the scene shown in the camera footage based on the scene information. In the reproduction processing, an abstracted image of a space shown in the camera footage is rendered and an abstracted image of the human is rendered thereon. In the reproduction processing, caption information generated from the linguistic information is added to the abstracted image of the space to generate a reproduced image.