Methods and apparatus to augment classification coverage for low prevalence samples through neighborhood labels proximity vectors

A two-stage classification process using a feature-based and appendix classifier with LSH Forests and custom distance metrics addresses the challenge of low prevalence malware samples, enhancing detection accuracy by identifying similar samples across prevalence levels.

US12645799B2Active Publication Date: 2026-06-02MCAFEE LLC

Patent Information

Authority / Receiving Office
US · United States
Patent Type
Patents(United States)
Current Assignee / Owner
MCAFEE LLC
Filing Date
2024-09-26
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing machine learning models struggle to accurately distinguish between noise and low prevalence malware samples, leading to gaps in detectability and poor generalization when classifying software as malicious or clean.

Method used

Implementing a two-stage classification process using a feature-based classifier and an appendix classifier that employs Locality Sensitive Hashing (LSH) Forests and custom distance metrics to identify similar samples, regardless of prevalence, thereby refining the classification of low prevalence malware samples.

Benefits of technology

Enhances the ability to classify low prevalence malware samples, closing the detectability gap and improving the overall accuracy of malware detection by leveraging neighborhood labels and proximity vectors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US12645799-D00000_ABST
    Figure US12645799-D00000_ABST
Patent Text Reader

Abstract

Disclosed examples include obtaining malicious neighbor samples based on a non-classified sample; obtaining clean neighbor samples based on the non-classified sample; generating a malicious proximity vector representing first distances between the non-classified sample and a first malicious neighbor sample from the malicious neighbor samples; generating a clean proximity vector representing second distances between the non-classified sample and a first clean neighbor sample from the clean neighbor samples; and classifying the non-classified sample as a clean sample or a malicious sample based on at least one of the malicious proximity vector or the clean proximity vector.
Need to check novelty before this filing date? Find Prior Art