File Vectorization and Machine Learning for Confidentiality Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual and rule-based systems for data classification are prone to inconsistencies, inaccuracies, and high costs due to the need for constant updates, limiting their ability to effectively classify and secure sensitive files based on confidentiality policies.

Innovation Solution

A system utilizing machine learning techniques, specifically neural networks and file vectorization, to automatically classify files based on confidentiality characteristics, optimizing storage strategies and enforcing security policies consistently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual classification methods are used, then classification accuracy can be maintained through human judgment, but the process becomes time-consuming and costly

Engineering Contradiction:
Improveclassification accuracyVSAvoidclassification time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary vectorization of files into structured representations that capture essential confidentiality characteristics. This preprocessing step enables subsequent rapid classification by the machine learning model without requiring time-consuming manual review of entire documents, thus reducing classification time while maintaining accuracy through pre-extracted features.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual human classification (mechanical system) with an automated machine learning system that uses neural networks to classify files based on vectorized representations. This substitution eliminates human time constraints and costs while maintaining or improving classification accuracy through consistent application of learned patterns across large volumes of files.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Extent of automation

If rule-based classification systems are used, then classification can be automated, but the systems require constant updates to maintain accuracy

Engineering Contradiction:
Improveautomation levelVSAvoidsystem maintenance complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The classification system transitions from static rule-based logic to dynamic machine learning models that automatically adapt to changing data patterns. The neural network continuously learns from new examples and updates its internal parameters, enabling the system to maintain high classification accuracy without manual rule updates, thus reducing maintenance complexity while preserving automation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The machine learning model performs self-updates by learning from new training data and automatically adjusting its classification thresholds and feature weights. This self-service capability eliminates the need for external experts to constantly update classification rules, reducing system maintenance complexity while maintaining automated operation and accuracy.

Inventive Principle:
Principle #25Self-service

3Reliability

If comprehensive file analysis is performed to ensure accurate classification, then classification reliability improves, but processing speed decreases

Engineering Contradiction:
Improveclassification reliabilityVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system extracts only the most relevant features from files during vectorization, creating condensed representations that capture essential confidentiality characteristics without including all raw data. This selective extraction maintains classification reliability by focusing on discriminative features while significantly reducing processing time compared to analyzing entire files comprehensively.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The classification process is segmented into distinct stages: vectorization of files into structured representations, feature extraction to identify key characteristics, and final classification by the machine learning model. This segmentation allows each stage to be optimized independently, maintaining reliability through thorough analysis while improving overall processing speed by eliminating redundant operations in each segment.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10354187B2Confidentiality of files using file vectorization and machine learning
Publication Date: 2019.07.16 HEWLETT PACKARD ENTERPRISE DEV LP
  • US10354187B2 patent drawing
  • US10354187B2 patent drawing
  • US10354187B2 patent drawing

AI summary

A method for confidentiality classification of files includes vectorizing a file to reduce the file to a single structured representation; and analyzing the single structured representation with a machine learning engine that generates a confidentiality classification for the file based on previous training. A system for confidentiality classification of files includes a file vectorization engine to vectorize a file to reduce the file to a single structured representation; and a machine learning engine to receive the single structured representation of the file and generate a confidentiality classification for the file based on previous training.