Selective Media Redaction via ML Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in effectively identifying and documenting sensitive information within media content while redacting it for specific applications, as existing methods often require generating and managing multiple versions of media content, leading to inefficient resource use and potential data exposure risks.
Innovation Solution
A method utilizing multilabel classification machine-learning models to segment and process audio or video content, generating metadata to identify sensitive information, and selectively redacting it based on context and individual associations, allowing for precise control over what is revealed in different applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple versions of media content are generated for different applications, then the need for selective redaction is met, but resource utilization efficiency deteriorates and data exposure risks increase
Solution Approach 1:
The media content is segmented into multiple audio content segments, and each segment is processed independently through the multilabel classification machine-learning model. This allows selective redaction of only the necessary portions containing sensitive information, rather than creating entirely separate versions of the complete media content, thereby improving resource utilization efficiency.
Solution Approach 2:
The system performs preliminary processing by segmenting the audio content and generating feature representations before redaction is needed. The multilabel classification model identifies sensitive information in advance, enabling efficient selective redaction when required without needing to maintain multiple pre-generated versions of the complete media content.
2Adaptability or versatility
If multiple versions of media content are generated for different applications, then the need for selective redaction is met, but data exposure risks worsen
Solution Approach 1:
By segmenting the audio content into smaller segments and processing them independently, the system minimizes the amount of data that needs to be redacted and stored in multiple versions. Only the specific segments containing sensitive information are flagged and redacted as needed, reducing the overall data exposure surface area compared to maintaining multiple complete versions.
Solution Approach 2:
The system extracts and identifies only the specific portions of audio content that contain sensitive information using the multilabel classification model. This extraction approach allows the system to maintain the original complete media content while selectively removing or redacting only the identified sensitive segments when required, rather than creating multiple complete versions.
3Ease of operation
If manual redaction processes are used, then control over redaction is maintained, but processing time and labor requirements increase
Solution Approach 1:
The system performs self-service by automatically processing the audio content through the multilabel classification machine-learning model to identify sensitive information and generate feature representations. This automated processing eliminates the need for manual analysis and redaction, significantly reducing processing time and labor requirements while maintaining consistent and controlled rediction through the AI model.
Solution Approach 2:
The patent replaces manual mechanical redaction processes with an automated machine-learning-based system. The multilabel classification model automatically analyzes audio segments, identifies sensitive information, and generates redaction instructions, substituting human labor with an automated intelligent system that is both faster and consistently controllable.
Data Source
AI summary
In general, various aspects of the present invention provide methods, apparatuses, systems, computing devices, computing entities, and/or the like for identifying and documenting certain subject matter found in media content, as well as selectively redacting the subject matter found in media content. In accordance with various aspects, a method is provided that comprises: obtaining first metadata for media content, the metadata identifying a context and a portion of content for an individual; identifying, based on the context, a certain subject matter for the media content; determining that a particular item associated with the subject matter is present in the portion of content associated with the individual; generating second metadata to document the particular item present in the portion of content; and using the second metadata to selectively redacting the item from the portion of content associated with the individual upon request to do so with respect to the individual.


