Video Sensitivity Classification Using Labeled Transcripts and AI
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for classifying and controlling the transmission of sensitive files, such as videos, within an enterprise network are inadequate due to a lack of accurate sensitivity determination based on the environment's data context.
Innovation Solution
A method utilizing machine learning models and generative AI to analyze videos, identify individuals, generate transcripts, and determine sensitivity classifications, with fine-tuned generative AI models and databases for prompt selection, enabling precise classification and control of file transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional file classification methods are used, then the system is simple to operate, but the sensitivity classification accuracy is insufficient
Solution Approach 1:
The system segments the video analysis process into multiple specialized components: a machine learning model for individual recognition, a speech-to-text module for transcript generation, a generative AI model for summary creation, and a classification model for sensitivity determination. Each component handles a specific aspect of the analysis, improving overall classification accuracy while maintaining manageable system complexity through modular design.
Solution Approach 2:
The patent introduces intermediate processing steps between the raw video input and the final sensitivity classification. A transcript module converts speech to text, a generative AI model creates contextual summaries, and these intermediates provide richer information to the classification model, thereby improving accuracy without directly increasing the complexity of the core classification function.
2Measurement precision
If comprehensive video analysis is performed to improve classification accuracy, then the sensitivity determination becomes more precise, but the processing time increases
Solution Approach 1:
The system performs preliminary actions by generating transcripts and summaries before the final sensitivity classification. These pre-processed elements capture essential information from the video content, allowing the classification model to make accurate determinations more efficiently without re-analyzing the entire video data during the classification stage.
Solution Approach 2:
The patent extracts key information from videos through specialized modules that pull out speech transcripts, individual identifications, and contextual summaries. This extraction process separates essential information from the raw video data, enabling faster and more accurate sensitivity classification by focusing only on the extracted relevant features rather than processing the complete video content.
3Reliability
If multiple AI models are deployed for detailed analysis, then the classification reliability improves, but the computational resources required increase
Solution Approach 1:
The system divides the analysis task across multiple specialized AI models, each optimized for a specific function: individual recognition, speech transcription, summary generation, and sensitivity classification. This segmentation allows each model to operate at optimal efficiency for its specific task, improving overall reliability while managing computational resource consumption through specialized rather than general-purpose processing.
Data Source
AI summary
A method and system for classifying a video file within an environment in which the file is located and when a file is classified as sensitive, controlling transmission of the file outside the environment. Classifying the video comprises analysing, using at least one machine learning model, the video to recognise any individuals in the video; obtaining a transcript of any speech in the video and generating, using the analysis, obtained transcript and a database of individuals linked to the environment, a labelled transcript which identifies each individual linked to the environment that is in the video. Information about each identified individual may be obtained from a connected database. A first generative AI model generates a text-based summary of the video by using the labelled transcript and information about identified individuals as prompts. A second generative AI model then determines a sensitivity classification of the video using the generated text-based summary.


