Feature selection using feature distribution in different subsets for machine learning in computer security applications

By partitioning data samples into training corpora based on device and user characteristics and selecting features with consistent frequency distributions, the system addresses the challenge of heterogeneous device environments, enhancing the reliability and efficiency of AI-based threat detection.

HK40135107APending Publication Date: 2026-07-17BITDEFENDER IPR MANAGEMENT

Patent Information

Authority / Receiving Office
HK · HK
Patent Type
Applications
Current Assignee / Owner
BITDEFENDER IPR MANAGEMENT
Filing Date
2026-05-06
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing AI-based computer security systems face challenges in selecting input features that are universally applicable across heterogeneous devices and effectively detect malware and online fraud, as they often require large training corpora and are prone to evasion by sophisticated malware.

Method used

A computer system and method that partitions data samples from multiple computing devices into training corpora based on device location, user identity, device type, or software configuration, and selects a reduced feature subset by comparing frequency distributions across corpora to train a threat detector.

Benefits of technology

This approach enhances the robustness and reliability of threat detection by focusing on features with consistent frequency distributions across diverse environments, reducing the need for large training datasets and improving detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A set of various data samples collected for a computer security application is divided into a plurality of training corpora according to criteria such as the identity of a data source. An initial set of features for characterizing the collected data samples is reduced to an optimal subset. For each candidate feature, a frequency distribution of feature values is determined on each of the training corpora. The feature selection process tends to features that are relatively similar across multiple corpora in their frequency distribution. A detector module is then trained to detect computer security threats from the reduced set of features.
Need to check novelty before this filing date? Find Prior Art

Description

(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202480058468.8 (22) Application Date 2024.09.11 (30) Priority Data 63 / 582,278 2023.09.13 US 18 / 401,364 2023.12.30 US (85) PCT International Application Entering National Phase Date 2026.03.12 (86) PCT International Application Application Data PCT / EP2024 / 075366 2024.09.11 (87) PCT International Application Publication Data WO2025 / 056605 EN 2025.03.20 (71) Applicant: Bitvane Intellectual Property Management Ltd. Address: Nicosia, Cyprus (72) Inventors: S. Smeu, E. Burchanu, E. Haller (74) Patent Agency: Beijing Law Alliance Intellectual Property Agency Co., Ltd. 11287 Patent Attorney: Zhang Shijun (51) Int.Cl. G06F 21 / 55 (2006.01) G06N 20 / 10 (2006.01) G06F 18 / 2115 (2006.01) (54) Invention Title: Feature Selection for Machine Learning in Computer Security Applications Using Feature Distributions in Different Subsets (57) Abstract: Based on criteria such as the identity of the data source, a collection of various data samples collected for computer security applications is divided into multiple training corpora. An initial feature set used to characterize the collected data samples is reduced to an optimal subset. For each candidate feature, the frequency distribution of the feature value is determined on each of the training corpora. The feature selection process tends to favor features whose frequency distribution is relatively similar across multiple corpora. The detector module is then trained to detect computer security threats based on the reduced feature set.Claims (3 pages), Description (13 pages), Drawings (7 pages), CN 121816573 A, 2026.04.07, CN 1 21 81 65 73 A. 1. A computer system comprising at least one hardware processor configured to: select a reduced subset of features from a plurality of features available for characterizing a data sample, wherein selecting the reduced subset of features comprises: partitioning a set of data samples acquired from a plurality of computing devices into a plurality of training corpora; selecting candidate features from the plurality of features; determining a first frequency distribution of feature values ​​of the candidate features over members of a first training corpus in the plurality of training corpora; determining a second frequency distribution of feature values ​​of the candidate features over a second training corpus in the plurality of training corpora; and determining whether to include the candidate features in the reduced subset of features based on the similarity between the first frequency distribution and the second frequency distribution; and, in response to selecting the reduced subset of features, training a threat detector to determine whether a target data sample indicates a computer security threat based on the reduced subset of features. 2. The computer system of claim 1, wherein the at least one hardware processor is configured to partition the set of data samples into the plurality of training corpora based on the location of the computing device providing each corresponding data sample. 3. The computer system of claim 1, wherein the at least one hardware processor is configured to partition the set of data samples into the plurality of training corpora based on the identity of the user of the computing device providing each corresponding data sample. 4. The computer system of claim 1, wherein the at least one hardware processor is configured to partition the set of data samples into the plurality of training corpora based on the acquisition time of each corresponding data sample. 5. The computer system of claim 1, wherein the at least one hardware processor is configured to partition the set of data samples into the plurality of training corpora based on the device type of the computing device providing each corresponding data sample. 6. The computer system of claim 1, wherein the plurality of computing devices are distributed among multiple company owners, and wherein the at least one hardware processor is configured to partition the set of data samples into the plurality of training corpora based on the owner of the computing device providing each corresponding data sample. 7. The computer system of claim 1, wherein the at least one hardware processor is configured to partition the set of data samples into the plurality of training corpora according to a software configuration file of the computing device providing each corresponding data sample, the software configuration file comprising a set of computer programs installed for execution on the corresponding computing device.8. The computer system of claim 1, wherein determining whether to include the candidate feature in the reduced feature subset further comprises: determining a plurality of similarity measurements, each of the plurality of similarity measurements quantifying the similarity between a pair of frequency distributions of the values ​​of the candidate feature; evaluating each of the pair of probability measurements on different corpora in the plurality of training corpora; and selecting the candidate feature in the reduced feature subset based on the average of the plurality of similarity measurements. 9. The computer system of claim 8, wherein the at least one hardware processor is configured to further select the candidate feature in the reduced feature subset based on the dispersion of the plurality of similarity measurements. Claims 1 / 3 Page 2 CN 121816573 A 10. The computer system of claim 1, wherein determining whether to include the candidate feature in the reduced feature subset further comprises: for each of the plurality of features, evaluating a feature-specific frequency distribution of the feature values ​​of the corresponding feature on the first training corpus; ranking the plurality of features based on the evaluated feature-specific frequency distribution; and selecting the candidate feature in the reduced feature subset based on the ranking result. 11. The computer system of claim 1, wherein the plurality of features are automatically constructed by a machine learning process, the machine learning process including training another threat detector to identify members of a set of data samples indicating the computer security threat. 12. A computer security method comprising employing at least one hardware processor of a computer system to: select a reduced subset of features from a plurality of features available for characterizing data samples, wherein selecting the reduced subset of features includes: partitioning a set of data samples acquired from a plurality of computing devices into a plurality of training corpora; selecting candidate features from the plurality of features; determining a first frequency distribution of feature values ​​of the candidate features over members of a first training corpus in the plurality of training corpora; determining a second frequency distribution of feature values ​​of the candidate features over a second training corpus in the plurality of training corpora; and determining whether to include the candidate feature in the reduced subset of features based on the similarity between the first frequency distribution and the second frequency distribution; and, in response to selecting the reduced subset of features, training a threat detector to determine whether a target data sample indicates a computer security threat based on the reduced subset of features. 13. The method of claim 12, further comprising dividing the set of data samples into the plurality of training corpora based on the location of the computing device providing each corresponding data sample.14. The method of claim 12, further comprising partitioning the set of data samples into the plurality of training corpora based on the identity of the user of the computing device providing each corresponding data sample. 15. The method of claim 12, further comprising partitioning the set of data samples into the plurality of training corpora based on the acquisition time of each corresponding data sample. 16. The method of claim 12, further comprising partitioning the set of data samples into the plurality of training corpora based on the device type of the computing device providing each corresponding data sample. 17. The method of claim 12, further comprising partitioning the set of data samples into the plurality of training corpora based on a software configuration file of the computing device providing each corresponding data sample, the software configuration file including a set of computer programs installed for execution on the corresponding computing device. 18. The method of claim 12, wherein the plurality of computing devices are partitioned among a plurality of company owners, the method comprising partitioning the set of data samples into the plurality of training corpora based on the owner of the computing device providing each corresponding data sample. 19. The method of claim 12, wherein determining whether to include the candidate feature in the reduced feature subset further comprises: determining a plurality of similarity measurements, each of the plurality of similarity measurements quantifying the similarity between a pair of frequency distributions of the candidate feature's values; evaluating each of the pair of probability measurements on different corpora in the plurality of training corpora; and selecting the candidate feature in the reduced feature subset based on the average of the plurality of similarity measurements. 20. The method of claim 19, further comprising selecting the candidate feature in the reduced feature subset based on the dispersion of the plurality of similarity measurements. 21. The method of claim 12, wherein determining whether to include the candidate feature in the reduced feature subset further comprises: for each of the plurality of features, evaluating a feature-specific frequency distribution of the corresponding feature's value on the first training corpus; ranking the plurality of features based on the evaluated feature-specific frequency distribution; and selecting the candidate feature in the reduced feature subset based on the ranking result. 22. The method of claim 12, wherein the plurality of features are automatically constructed by a machine learning process, the machine learning process including training another threat detector to identify members of a set of data samples indicating the computer security threat.23. A non-transitory computer-readable medium storing instructions, when executed by at least one hardware processor of a computer system, the instructions causing the computer system to: select a reduced subset of features from a plurality of features available for characterizing data samples, wherein selecting the reduced subset of features includes: partitioning a set of data samples acquired from a plurality of computing devices into a plurality of training corpora; selecting candidate features from the plurality of features; determining a first frequency distribution of feature values ​​of the candidate features over members of a first training corpus in the plurality of training corpora; determining a second frequency distribution of feature values ​​of the candidate features over a second training corpus in the plurality of training corpora; and determining whether to include the candidate features in the reduced subset of features based on the similarity between the first frequency distribution and the second frequency distribution; and, in response to selecting the reduced subset of features, training a threat detector to determine whether a target data sample indicates a computer security threat based on the reduced subset of features. Claims 3 / 3 Page 4 CN 121816573 A Feature selection for machine learning in computer security applications using feature distributions in different subsets

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 582,278, filed September 13, 2023, entitled “Feature Selection for Robust Anomaly Detectors,” the entire contents of which are incorporated herein by reference. Background Art

[0003] The present invention relates to machine learning, and more particularly, to the optimal selection of input features for classifiers used in computer security applications, including the detection of malware, intrusions, and online fraud.

[0004] Computer security is an important branch of information technology aimed at protecting users and computers from malware, intrusions, and fraudulent use. Malware, in its various forms (such as computer viruses, spyware, and ransomware), affects millions of devices, making them vulnerable to fraud, loss of data and sensitive information, identity theft, and loss of productivity.

[0005] Another persistent threat comes from online fraud, particularly in the form of phishing and identity theft. Sensitive identity information (such as usernames, IDs, passwords, social security and medical records, bank and credit card details) obtained through fraudulent means by international criminal networks operating on the Internet is used to withdraw private funds and / or is further sold to third parties.In addition to the direct economic losses to individuals, online fraud has a range of negative economic consequences, such as increased security costs for companies, higher retail prices and bank fees, declining stock values, lower wages, and reduced tax revenue.

[0006] The explosive growth of mobile computing only exacerbates computer security risks, as millions of smartphones and tablets are constantly connected to the Internet and become potential targets for malware and fraudulent attempts.

[0007] Various computer security methods and software are available to protect users and computers from such threats. Modern systems and methods increasingly rely on artificial intelligence (AI) pre-trained to distinguish between malicious and benign samples. A typical example of an AI-based malware detector involves a neural network configured to receive a vector of feature values ​​representing an input sample and produce an output indicating whether the corresponding sample is malicious.

[0008] However, AI-based methods also face significant technical challenges of their own. One example is the selection of input features. Typically, there are no clear or universal guidelines for selecting which features of target software are more likely to reveal malice and / or distinguish between malicious and benign behavior. Furthermore, a single set of malware signatures cannot reliably work across highly heterogeneous sets of devices (e.g., desktop computers, mobile computing platforms (smartphones, wearables, etc.), and Internet of Things (IoT) devices). The problem is further complicated by sophisticated malware deliberately attempting to evade detection. Malware may tailor its behavior based on device type (e.g., smartphones vs. tablets, one manufacturer or model vs. another), operating system type, the current geographic location of the corresponding device, etc. Some malware further selects victims by searching for indicators of a user's value to the attacker on the corresponding device. For example, malware may determine which other software is currently installed on the corresponding device and search for specific applications, such as banking, social media, etc. Other malware may monitor patterns of user access to various applications, online resources, etc. Such malware may then launch attacks when attacks on only carefully selected devices and / or carefully selected users are deemed more likely to be rewarding. Specification 1 / 13 pages 5 CN 121816573 A

[0009] To address the variability of malware and the heterogeneity of host devices, some conventional methods have significantly increased the count of features and thus increased the size of AI models in an attempt to improve their performance. However, the high cost of implementing and training large neural networks is well-known in the industry, and they typically require large training corpora that are difficult to obtain, annotate, and maintain. Another common approach uses unsupervised training, in which the AI ​​system is configured to build its own feature set based on the available training corpus.However, such self-generated features generally have no informational value to human users and cannot be guaranteed to perform as expected when applied to data samples that the corresponding AI system has not previously seen.

[0010] For all the above reasons, there is great interest in developing more robust and reliable AI-based computer security systems and methods. Summary of the Invention

[0011] According to one aspect, a computer system includes at least one hardware processor configured to select a reduced subset of features from a plurality of features that can be used to characterize data samples. Selecting the reduced subset of features includes partitioning a set of data samples obtained from a plurality of computing devices into a plurality of training corpora, selecting candidate features from the plurality of features, determining a first frequency distribution of the feature values ​​of the candidate features over members of a first training corpus in the plurality of training corpora, and determining a second frequency distribution of the feature values ​​of the candidate features over a second training corpus in the plurality of training corpora. Selecting the reduced subset of features further includes determining whether to include the candidate features in the reduced subset based on the similarity between the first frequency distribution and the second frequency distribution. The at least one hardware processor is further configured to train a threat detector in response to selecting the reduced feature subset, to determine whether a target data sample indicates a computer security threat based on the reduced feature subset.

[0012] According to another aspect, a computer security method includes employing at least one hardware processor of a computer system to select a reduced feature subset from a plurality of features that can be used to characterize a data sample. Selecting the reduced feature subset includes partitioning a set of data samples obtained from a plurality of computing devices into a plurality of training corpora, selecting candidate features from the plurality of features, determining a first frequency distribution of the feature values ​​of the candidate features over members of a first training corpus in the plurality of training corpora, and determining a second frequency distribution of the feature values ​​of the candidate features over a second training corpus in the plurality of training corpora. Selecting the reduced feature subset further includes determining whether to include the candidate feature in the reduced feature subset based on the similarity between the first frequency distribution and the second frequency distribution. The method further includes employing the at least one hardware processor to train a threat detector in response to selecting the reduced feature subset, to determine whether a target data sample indicates a computer security threat based on the reduced feature subset.

[0013] According to another aspect, a non-transitory computer-readable media storage instruction, when executed by at least one hardware processor of a computer system, causes the computer system to select a reduced subset of features from a plurality of features that can be used to characterize a data sample.Selecting the reduced feature subset includes dividing a set of data samples obtained from multiple computing devices into multiple training corpora, selecting candidate features from the multiple features, determining a first frequency distribution of the feature values ​​of the candidate features on members of a first training corpus in the multiple training corpora, and determining a second frequency distribution of the feature values ​​of the candidate features on a second training corpus in the multiple training corpora. Selecting the reduced feature subset further includes determining whether to include the candidate features in the reduced feature subset based on the similarity between the first frequency distribution and the second frequency distribution. The instructions further cause the computer system to train a threat detector in response to selecting the reduced feature subset to determine whether a target data sample indicates a computer security threat based on the reduced feature subset. Specification 2 / 13 pages 6 CN 121816573 A Brief Description of the Drawings

[0014] The foregoing aspects and advantages of the invention will become better understood by reading the following detailed description and referring to the accompanying drawings, in which:

[0015] FIG1 illustrates multiple client devices protected from computer security threats according to some embodiments of the invention.

[0016] FIG2 illustrates an exemplary artificial intelligence (AI) training system for collecting data from multiple data sources according to some embodiments of the present invention.

[0017] FIG3 shows exemplary components of an AI training system according to some embodiments of the present invention.

[0018] FIG4 illustrates exemplary components of a threat detector module according to some embodiments of the present invention.

[0019] FIG5 shows an exemplary general feature vector according to some embodiments of the present invention.

[0020] FIG6 shows an exemplary sequence of steps performed by an AI training system according to some embodiments of the present invention.

[0021] FIG7 shows an exemplary feature selection process according to some embodiments of the present invention.

[0022] FIG8 shows an exemplary statistical distribution of the values ​​of selected data features according to some embodiments of the present invention.

[0023] FIG9 shows an exemplary hardware configuration of a computer system programmed to perform some of the methods described herein. Detailed Description

[0024] In the following description, it should be understood that all stated connections between structures may be direct operational connections or indirect operational connections through intermediate structures. A set of elements comprises one or more elements. Any statement about an element is understood to refer to at least one element. Multiple elements comprise at least two elements. Unless otherwise required, any described method steps need not be performed in the specific order stated. The first element derived from the second element (e.g., data) encompasses the first element equal to the second element, as well as the first element generated by processing the second element and optionally other data. Making a decision based on parameters encompasses making a decision based on parameters and optionally other data.Unless otherwise specified, some quantity / data indicators may be the quantity / data itself, or indicators different from the quantity / data itself. A computer program is a sequence of processor instructions that performs a task. The computer program described in some embodiments of the present invention may be a standalone software entity or a sub-entity of another computer program (e.g., subroutines, libraries). The term 'database' is used herein to refer to any organized collection of data. Computer-readable media encompasses non-transitory media, such as magnetic, optical, and semiconductor storage media (e.g., hard disk drives, optical disks, flash memory, DRAM), and communication links, such as conductive cables and fiber optic links. According to some embodiments, the present invention particularly provides a computer system comprising hardware (e.g., one or more processors) programmed to perform the methods described herein and computer-readable media encoded instructions for performing the methods described herein.

[0025] FIG1 illustrates a plurality of client devices 12a to d protected from computer security threats according to some embodiments of the present invention. Exemplary client devices 12a to d include personal computer systems, enterprise mainframes, mobile computing platforms (e.g., laptops, tablets, smartphones), entertainment devices (e.g., TVs, game consoles), wearable devices (e.g., smartwatches, fitness bands), home appliances (e.g., thermostats, refrigerators), and any other electronic devices, which include hardware processors, memory, and communication interfaces that enable the respective devices to communicate with other devices / computer systems. In some embodiments, each client device 12a to d includes a threat detector module 40 configured to protect the respective device from computer security threats. Any such detector module may be embodied as a set of linked computer programs executing on at least one hardware processor of the respective client device. The operation of the threat detector is described in more detail below. Specification 3 / 13 pages 7 CN 121816573 A

[0026] Exemplary client devices 12a to d are connected to a communication network 15, which may include a local area network (e.g., home network, enterprise network, etc.), a wide area network, and / or the Internet. Network 15 typically represents a set of hardware (physical layer) and software interfaces that enable data transfer between devices 12a to d and other entities connected to network 15.

[0027] Figure 1 further illustrates a security server 14 connected to communication network 15. Server 14 typically represents a group of communication-coupled computer systems that may or may not be physically close to each other. Server 14 can be configured to protect multiple clients 12a to d from threats by analyzing data received from the respective devices. Such protection can be customized for each device, group of devices, user, or user group, and various parameters of the corresponding protection service can be defined via subscription / service agreements, etc.

[0028] In some embodiments, the security server 14 may collect various data from individual client devices without requiring the active participation of the respective devices. For example, some embodiments extract forensic data by intercepting network traffic at various points within the communication network 15. In alternative embodiments, each client device 12a to d may cooperate with the server 14 to protect each respective device. In other words, computer security activities may be divided between components executing on the respective device and components executing on the server 14. For example, each client device may collect and preprocess data indicating the behavior of software executing on the respective device, and then send such data to the server 14.

[0029] Subsequently, the server 14 may run an example of a threat detector module 40, which is configured to analyze data received or collected from each client and return a security indicator to the respective client, the security indicator indicating whether the analyzed data indicates a computer security threat.

[0030] FIG1 further illustrates an artificial intelligence (AI) training system 16 communicationally coupled to the security server 14. AI training system 16 typically represents a set of interconnected computer systems tasked with training detector modules, including artificial intelligence components, to detect computer security threats based on data received from client devices.

[0031] Training as used herein represents a machine learning process by which a set of parameters of a threat detector module are adjusted to improve the performance of the respective detector. In some exemplary embodiments described below, the various components of the threat detector may comprise an artificial neural network with a large number of adjustable parameters (e.g., synaptic weights, etc.). In such embodiments, training involves adjusting the respective parameters to minimize an objective function commonly referred to in the field as cost. An exemplary cost function may be determined based on the difference between the output of the respective detector when a training sample is presented and the expected or desired output associated with the respective training sample. Several training processes are known in the field, including various versions of supervised, semi-supervised, and unsupervised learning processes.

[0032] Figure 2 illustrates AI training system 16 receiving training data from multiple data sources 13a to e. Such sources may include personal computers and / or computing devices (data sources 13d to e) and computer systems comprising multiple individual machines (e.g., sources 13a to c). In one such example, data source 13a represents a home where multiple personal devices (e.g., smartphones, tablets, home appliances, entertainment devices, wearable devices, etc.) access the Internet via a gateway device / edge router and contribute data to the AI ​​training system 16. Meanwhile, exemplary data sources 13b to c may represent different enterprise computer networks or different network domains of an enterprise network.

[0033] Figure 3 illustrates exemplary components of the AI ​​training system 16 according to some embodiments of the present invention.Some or all of the components described may be embodied as computer programs executing on the hardware processor of system 16. In alternative embodiments, some components may be embodied in dedicated hardware, such as ASICs, FPGAs, etc., or as a combination of hardware and software. AI training system 16 includes a data selector component 22 configured to receive data samples 18a to c from corresponding data sources 13a to c, and to construct multiple training corpora 20a to d based on the received data samples. Data samples 18a to c generally represent any data type known in the field to be of informational value to computer security. The content of such samples depends on the type of computer security threat described in the specification 4 / 13 page 8 CN 121816573 A (e.g., malware and online fraud, etc.). Some instances contain encodings of sequences of computational events occurring on the corresponding computer (e.g., the contents of system logs), and the contents of sectors of the corresponding computer's memory. Other exemplary data samples 18a to c contain indicators of network traffic, such as various service metadata. Another exemplary data sample 18a to c contains the content of electronic messages (e.g., messages exchanged via instant messaging applications, email messages, social media posts, SMS messages, etc.). Other exemplary samples 18a to c may contain media files (sound, images, videos, etc.).

[0034] In some embodiments, training corpora 20a to d contain data samples 18a to c, which are selectively organized into individual corpora according to various criteria (e.g., the identity of the computer / data source, the identity of the user, time, geographic location and / or the network address of the corresponding computer / data source, etc.). More details on organizing training data into individual corpora are given below.

[0035] In some embodiments, the AI ​​training system 16 further includes a training engine 24 connected to an example of the feature selector 30 and the detector module 40. The engine 24 is configured to manage and coordinate the activities of the detector module 40 and the feature selector 30, and to implement the process for training the detector 40. Engine 24 is further configured to output the results of training, such as detector specification 26, which includes a set of optimal values ​​for detector parameters (e.g., network architecture specification, synaptic weights, etc.) generated by training.

[0036] Figure 4 shows an exemplary structure of threat detector 40, which includes a feature extractor 42 connected to classifier 44. In some embodiments, detector 40 is configured to receive data sample 18 and output a security label 28 indicating whether sample 18 indicates a threat. Training examples of threat detector 40, which is performed as part of AI training system 16, receive input samples 18a to c selected by training engine 24 from training corpora 20a to d, as described in further detail below.

[0037] The types of threats detected by module 40 vary depending on the embodiment.Without loss of generality, most of the following description will focus on embodiments that detect the presence or activity of malware. In this case, security label 28 may indicate whether the corresponding client device is infected. However, those skilled in the art will appreciate that the described embodiments are suitable for detecting other threats, such as online fraud, intrusion / hacking, deepfakes, etc. In yet another class of applications, threat detector 40 includes an anomaly detector configured to determine whether the corresponding data sample 18 follows a normal or expected behavior pattern of the corresponding client device or deviates from such a normal behavior pattern. In such embodiments, security label 28 may indicate whether the data sample 18 is anomalous. Anomalies may further indicate a threat, such as intrusion / hacking.

[0038] Feature extractor 42 is configured to determine a set of feature values ​​that commonly characterize the data sample 18. Exemplary feature vector 32 is illustrated in FIG. 5 and includes feature values ​​f1…fN, each feature value fi including the value of the corresponding feature Fi of sample 18. Feature Fi herein refers generally to any attribute of data sample 18 known in the field of computer security that can be used to determine whether a client device is subject to a computer security threat, such as malware, fraud, hacking, etc. Feature Fi can be determined or selected "manually" (i.e., by a human operator) or automatically, such as through unsupervised learning. Some of these features may indicate a threat on their own, while others may only indicate the presence of a threat when they occur simultaneously with other features. Exemplary features include so-called 'static' features, such as features extracted from the contents of storage, files or folders (e.g., the Android® manifest file), electronic messages, or images displayed on the screen of the corresponding client device. Other static features can be determined based on current OS settings, the contents of the OS registry, etc. Other exemplary features include behavioral or so-called 'dynamic' features, which indicate whether a particular software entity has performed a specific action, such as opening a file, changing access permissions, initiating a child process, injecting code into another software entity, sending electronic communications (e.g., HTTP requests to remote resources), etc. Some feature values ​​fi can be derived from the values ​​of other features fj. Depending on the type of the corresponding feature, feature value fi can be a number, string, Boolean, etc. In a typical embodiment, the feature count N is relatively large, for example, hundreds or thousands, but this description generally applies to any number of features as described on page 5 / 13 of the specification, CN 121816573 A.

[0039] To determine the feature vector 32, the feature extractor 42 may implement any method known in the field of computer security. Evaluating static features may involve signature matching by hashing or other methods. Signature matching typically determines whether the memory of the client device stores a specific sequence of code / instructions commonly referred to in the field as a malware signature.Evaluating dynamic characteristics may involve parsing a series of computational events (e.g., system logs) to detect the occurrence of specific hardware or software events, such as application installation, uninstallation, and updates; process / application startup and termination; child process creation (e.g., forking); dynamic loading / unloading of libraries; execution of specific processor instructions (e.g., system calls); file events such as file creation, writing, and deletion; and setting various OS parameters (e.g., Windows® Registry events, permission / privilege changes). Other exemplary events detected include requests to access peripheral devices (e.g., hard drives, SD cards, network adapters, microphones, cameras); requests to access remote resources (e.g., Hypertext Transfer Protocol – HTTP requests to access a specific URL, intent to access a document repository over a local network); requests specified using a particular Uniform Resource Identifier scheme (e.g., mailto: or ftp: requests); and intents to send electronic messages (e.g., email, Short Message Service – SMS, etc.). Other exemplary events include moving a user interface / window displayed by the corresponding client device into and / or out of focus / foreground.

[0040] In some embodiments, extracting dynamic features may further include detecting various time-related events, such as inactive periods, i.e., time intervals between events and / or time intervals when the corresponding client device is idle, has no registered user activity, or is only performing internal system tasks. Such inactive periods may be further divided into short intervals (e.g., on the order of seconds) and long intervals (e.g., on the order of minutes to hours). Other time-related events may include, for example, a series of events occurring in a rapid, continuous / burst of activity.

[0041] Other exemplary dynamic features include receiving and / or displaying specific types of content, such as SMS messages containing hyperlinks, HTML documents containing login forms, payment interfaces, advertisements, etc.

[0042] Extracting dynamic features specific to mobile devices or particularly related to the security of mobile devices includes, for example, detecting events such as screen switching (on / off), changes in application labels / names / icons, and screen captures. Other examples include requests to grant specific types of permissions (e.g., administrator, accessibility), dynamic requests (i.e., permissions granted during various stages of execution, not at installation time), and granting persistence (e.g., foreground services dynamically started by the corresponding application). Still other examples include attempts to prevent the uninstallation of the corresponding application and displaying an overlay at the top of the OS settings interface (such overlays may trick unsuspecting users into granting unnecessary permissions to the corresponding application).

[0043] In some embodiments, feature Fi is a higher-order feature derived from other more fundamental features of the input data. For example, extractor 42 may evaluate a set of primary features of data sample 18, such as indicators of the occurrence of specific events on the corresponding client device.Then, the extractor 42 can combine multiple principal features into composite features, such as linear combinations of principal features. In one such instance, the extractor 42 performs singular value decomposition (SVD) or principal component analysis (PCA) on the evaluated principal features to produce a set of composite features, wherein at least one composite feature Fi includes principal components of the corresponding principal features.

[0044] In another exemplary embodiment of constructing composite features, the extractor 42 includes a set of artificial neural networks (collectively referred to in the field as encoders) configured to compute the projection of the input data sample 18 into an abstract vector space, which is generally referred to as an embedding. In other words, individual features Fi may include the projection of the data sample 18 along individual axes of the embedding space. Meanwhile, the elements fi of the feature vector 32 may include the coordinates of points in the embedding space, the corresponding coordinates fi being determined by a specific mathematical transformation based on the current values ​​of some principal attributes / features of the data sample 18. The corresponding embedding space / mathematical transformation may be automatically determined via a machine learning process, wherein some parameters of the feature extractor 42 are tuned to meet a predetermined objective, such as minimizing a cost function determined based on the input and output of the feature extractor 42 and / or the detector 40. In some embodiments, the extractor 42 and classifier 44 are trained together.

[0045] The classifier 44 includes modules (e.g., a set of computer programs) configured to determine the security label 28 based on the feature vector 32. For example, the classifier 44 may be configured to distinguish between multiple classes / categories of data and determine which class / category the data sample 18 falls into based on the feature vector 32. Exemplary classes include classes representing normal behavior and another class representing abnormal behavior. Other exemplary classes include clean and infected. Still other exemplary classes indicate different types of malicious proxies or attack strategies. Those skilled in the art will appreciate that such classes are given only as examples and are not intended to be limiting. Furthermore, the examples presented herein may be adapted as needed to detect other computer security threats, such as online fraud, deepfakes, etc.

[0046] The classifier 44 may implement any classification method known in the field of data mining. Examples include decision trees, clustering methods (e.g., k-means or correlation), and artificial neural networks, etc. The latter may further include feedforward networks, autoencoders, and various components of neural networks specifically designed for image, text, and natural language processing, such as recurrent neural networks (RNNs), long short-term memory (LSTM) architectures, convolutional neural networks, transformer neural networks (e.g., generative pre-trained transformers - GPT), etc. Some architectural and functional details of such neural networks are beyond the scope of this description.

[0047] Figure 6 illustrates an exemplary sequence of steps performed by the AI ​​training system 16 in some embodiments of the invention.Steps 202 to 204 collect training data, such as samples 18 received from various data sources 13a to e, until accumulation conditions are met. Various accumulation conditions can be used, such as collecting a predetermined total number of samples, a predetermined number of samples for each data source 13a to e, a predetermined number of samples for each device type, and collecting incoming samples for a predetermined amount of time. Then, in step 206, the data selector 22 can construct multiple training corpora (see, for example, corpora 20a to d in FIG3) by organizing the collected training data according to various criteria.

[0048] Some embodiments rely on the collected data being considered as observations comprising at least two components. The first component, referred to herein as a “signal,” carries information related to computer security threats, such as computational events indicating the presence or activity of malware. The second component, referred to herein as “background,” carries information related to the context or environment in which the malware operates. The background component, for example, corresponds to benign, normal activity performed on the corresponding computing device. However, signal and background components are intertwined within the collected data, and it is not prior to know which features of the data are more informative or better characterize one or the other of the components. Some embodiments utilize the assumption that attempting to separate the signal from the background may benefit from partitioning heterogeneous data sets into individual corpora 20a to d (Figure 3), where each individual corpus is characterized by a different type of background, and the signal components are similar across multiple corpora. Optimal feature selection can then include comparing the performance of selected features across different corpora 20a to d, as described below. Therefore, features that are relatively insensitive to switching between one corpus and another should be more suitable for characterizing the signal than other features.

[0049] Some embodiments partition the collected data samples 18a to c (Figure 3) into separate training corpora 20a to d based on various attributes of the data source providing each corresponding sample. An exemplary criterion is the location of the corresponding data source 13a to e. In other words, different samples 18a to c received from the same location can be placed into the same corpus, while different samples 18a to c received from different locations can be placed into different training corpora 20a to d. The location in this paper encompasses physical (physical address, geographic location, etc.) and logical / virtual locations (network domain, network address, Internet Protocol-IP address, Uniform Resource Identifier-URI, etc.).

[0050] The strategy of constructing different training corpora based on the location of the data source relies on the observation that each data source (e.g., each household, company, network domain, etc.) may have its own particularities in how it uses the corresponding computing devices. For example, an accounting firm's computer may be used differently than an engineering company's or university's computer. Furthermore, different levels of granularity can be used for partitioning into individual corpora.For example, some embodiments can distinguish data samples 18a to c received from different departments of the same organization. In one such example, data samples collected from the marketing and production departments of the same company are placed into different training corpora 20a to d.

[0051] In one example of creating a corpus based on geographic location, data samples 18a to c collected from devices / sources located in the same country or region can be grouped together into the same training corpus, while samples collected from different countries / regions can be merged into different corpora 20a to d. In an alternative embodiment, data sources / devices can be distinguished based on language, keyboard layout, computing location settings, etc., so that data samples received from devices with the same characteristics are grouped into the same training corpus.

[0052] Alternative criteria for constructing individual training corpora 20a to d include the identity of the user providing the computing device for the corresponding data samples 18a to c. In other words, different training corpora 20a to d can consist of data received from different users or user groups. When a user operates multiple computing devices, data samples collected from all such devices can be placed into the same training corpus. Conversely, when a device is used by multiple users, data samples 18a to c collected from the respective devices can be placed into different training corpora based on the identity of the user operating the respective device when collecting the corresponding data samples.

[0053] Another exemplary criterion for constructing training corpora 20a to d includes the device type of the source of the respective data samples. Device type can be defined by appliance type (e.g., smartphones and desktop computers and thermostats), manufacturer (e.g., Apple Inc. or Samsung Electronics), and communication protocol (Bluetooth™ and Wi-Fi™), etc. In some such embodiments, data samples 18a to c received from devices of the same type are grouped together in the same training corpus. Conversely, samples collected from devices of different types can be placed into different training corpora.

[0054] Yet another exemplary criterion for constructing training corpora 20a to d includes a software profile of the device supplying each data sample. In some embodiments, the software profile includes a specific set of software applications or application types installed and / or used on the respective computing device. Some such embodiments may install a profile agent on each device providing training data to the AI ​​training system 16. The profile agent's task is to compile a list of software applications currently installed and / or used on the respective device. The data selector 22 can then categorize contributing data sources according to the software profiles, such that data samples 18a to c received from devices with the same software profile are merged together into the same training corpus, while samples from devices with different software profiles are distributed into different training corpora.Such embodiments rely on observations that device / user behavior may inherently depend on the type of software installed on the respective machine. For example, the behavior of a gamer may differ significantly from that of an accountant, and the difference may be reflected in the type of software installed and / or frequently used on the respective device. Therefore, placing data samples collected from these categories of users into different training corpora ensures that the “noise” or “background” component of a corpus is different across different corpora.

[0055] Another exemplary criterion for classifying data samples 18a to c into training corpora 20a to d is time. For example, each training corpus 20a to d may consist of samples collected during different time intervals. Exemplary time intervals can distinguish between the two based on the observation that computational behavior may differ between working hours and leisure / weekends / holidays.

[0056] Yet another exemplary criterion for classifying training samples into different corpora (see step 206 in FIG. 6) includes a method for constructing data samples 18a to c. Such criteria are particularly applicable to synthetic training data specifically generated for training detector 40. Some examples of this include deepfake data (AI-generated or enhanced images, videos, audio, etc.) and machine-generated text (e.g., the output of various chatbots). In such examples, data selector 22 can distinguish samples generated using different methods. For example, samples 18a to c generated using the same method (or group of methods) can be placed in the same training corpus, while samples generated using different methods can be merged into different training corpora.

[0057] Some other alternative embodiments may use more sophisticated methods of organizing the collected training data, such as applying clustering algorithms to divide the collected data samples 18a to c into multiple clusters, where substantially similar samples belong to the same cluster, regardless of their data source / user. In such embodiments, each distinct sample cluster (or subset of clusters) may form separate training corpora 20a to d. One such example is applicable to the detection of online fraud, where a collection of messages can be divided into multiple clusters according to various criteria (e.g., according to the type or category of fraud (e.g., insurance fraud versus Nigerian fraud versus cryptocurrency fraud)). Different training corpora 20a to d can then be constructed to contain the content of the selected clusters.

[0058] In some embodiments where the feature extractor 42 includes AI components, for example, when the feature extractor 42 is configured to compute the embedding representation of the input data samples, the extractor 42 needs to be trained to produce the desired feature vectors (step 208 in FIG. 6).The cost function used in training is determined based on the output of classifier 44, and training includes tuning the parameters of both the feature extractor and the classifier according to the value of the corresponding cost function. Some feature extractors 42 are trained together with classifier 44. In an alternative embodiment, the feature extractors are trained independently of classifier 44 and / or using different training corpora.

[0059] Step 208 configures the feature extractor 42 to evaluate an initial feature set Fi for characterizing the collected data samples 18a to c. Then, step 210 performs a feature selection process that includes reducing the initial feature set to an optimal, reduced subset of features. Figure 7 illustrates an exemplary feature selection process according to some embodiments of the invention. The sequence of steps 220 to 222 iterates through all available initial features. For each feature Fi considered here as a candidate feature, steps 224 to 226 then iterate through all training corpora 20a to d constructed in step 206 (Figure 6). For each such training corpus Cj, step 228 may calculate the distribution of feature Fi on corpus Cj, i.e., calculate the value fi of the corresponding candidate feature Fi for at least a portion of the corresponding corpus of the samples. Some embodiments may further determine the frequency distribution of feature Fi by dividing the range of the corresponding feature into bins and counting how many data samples 18a to c within the corresponding corpus Cj are characterized by the fi value in each bin. The resulting count may be normalized to the total count of samples in the corresponding corpus. Figure 8 shows exemplary frequency distributions 50a to f, where the left column represents the frequency distribution of exemplary feature F1 and the right column represents the frequency distribution of another exemplary feature F2. The top, middle, and bottom rows show the distributions calculated on three different exemplary training corpora C1, C2, and C3, respectively. Each additional figure includes a set of bins (rectangles), where each bin corresponds to a different feature value fi (or value range), and the height of the corresponding bin / rectangle indicates the count of examples / frequencies of the corresponding feature value within the corresponding training corpus Cj.

[0060] When all frequency distributions of the selected features have been evaluated (step 224 returns No), in step 230, training engine 24 may compute a set of similarity measures, each such measure quantifying the similarity of the corresponding pair {j, k} of the frequency distributions of the feature Fi determined in step 228 on the training corpora Cj and Ck. Using the exemplary distributions illustrated in Figure 8, the similarity measures may quantify the similarity between distributions 50a and 50c, while the similarity measures may quantify the similarity between distributions 50d and 50f. Therefore, step 230 may return a set of feature-specific similarity measures, where M represents the counts of training corpora 20a to d.

[0061] Step 228 may use any method of calculating the similarity between two statistical distributions.An exemplary embodiment determines the measurement based on the Wasserstein metric:

[0062] [1]

[0063] where the range of feature Fi is divided into n bins, b indexes individual bins, and represents the frequency (e.g., normalized count) of the value fi in bin b that falls within the corpus Cj, and is a positive number (e.g., = 1). Alternatively, it can be determined based on the Kullback-Leibler divergence: Specification 9 / 13 pages 13 CN 121816573 A

[0064] [2]

[0065] When the corresponding statistical distributions are similar, the Wasserstein and Kullback-Leibler distances are relatively small, and otherwise relatively large.

[0066] When all similarity measures have been evaluated for all currently available features Fi and all training corpora (step 220 returns no), step 232 can rank features Fi based on the calculated similarity measures. The ranking Ri of feature Fi is determined based on a set of corresponding similarity measures, such as the average (e.g., mean, median) of set Si and / or the degree of dispersion (e.g., range, variance, standard deviation).

[0067] Then, a further step 234 may select a reduced subset of features based on the ranking determined in step 232. Some embodiments rely on the observation that features whose distribution is relatively stable (i.e., less variable) across multiple training corpora may be more robust to changes between data sources. In other words, selecting such features as input to classifier 44 may make classifier 44 more robust when applied to data from new sources (i.e., data it did not previously encounter in training). Thus, in some embodiments, a feature Fi whose average indicates that its frequency distribution is relatively stable across multiple corpora receives a relatively higher ranking than a feature whose frequency distribution changes significantly from one training corpus to another. For example, some embodiments may rank features based on an average similarity measure across all training corpus pairs, where a smaller average corresponds to a higher ranking. In the example of Figure 8, some embodiments rank feature F2 higher than feature F1 because the distribution of feature F1 changes less from corpus C1 to C2 and C3 than the distribution of feature F2. Then, step 234 may select a predetermined number or predetermined proportion (e.g., 80%, 50%, etc.) of the highest-ranked feature Fi into the reduced feature subset.

[0068] In response to the completion of the feature selection process, in step 212 (Figure 6), training engine 24 may reconfigure classifier 44 to specifically use the reduced set of selected features as input. In a further step 214, engine 24 may train the reconfigured classifier / threat detector on some or all training corpora 20a to d.For example, such training may include presenting training data samples to classifier 44 and adjusting the parameters of classifier 44 based on the output generated by classifier 44.

[0069] In some embodiments, step 216 determines whether the trained classifier meets predetermined performance criteria, such as the receiver operating characteristic (ROC) curve determined on the training corpus. Other performance criteria may include computational cost, speed of operation, etc. If classifier 44 is found to be satisfactory, then in step 218, AI training system 16 may output detector specification 26, which may include, for example, the architecture specification and internal parameter values ​​generated by training. Detector specification 26 may then be used by security server 14 and / or client devices 12a to d (FIG. 1) to instantiate a local version of the trained threat detector 40.

[0070] In some embodiments, when the evaluation shows that detector 44 does not meet the performance criteria (step 216 returns no), training engine 24 may rerun at least a portion of the feature selection process (step 210), for example, to select other features into a reduced feature subset, or to change the proportion of features included in the reduced feature subset. Computer experiments applying some of the methods described herein to image processing have found that the classifier's performance steadily improves as the initial feature set is gradually reduced to approximately 50% of its initial content, after which performance declines significantly. Some embodiments may use a similar gradual reduction of the feature set in computer security applications to monitor the classifier's performance and stop the feature selection process before the classifier's performance declines.

[0071] Figure 9 illustrates an exemplary hardware configuration of a computer system 80 programmed to perform some of the methods described herein. Computer system 80 generally represents any of the client devices 12a to d in Figure 1, as well as a security server 14 and AI training specification 10 / 13 pages 14 CN 121816573 A system 16. The illustrated computing system is a personal computer; other devices (e.g., servers, mobile phones, tablet computers, and wearable devices) may have slightly different configurations.

[0072] Processor 82 includes physical means (e.g., a microprocessor, a multi-core integrated circuit formed on a semiconductor substrate) configured to perform computations and / or logical operations on a set of signals and / or data. Such signals or data may be encoded and delivered to processor 82 in the form of processor instructions (e.g., machine code).

[0073] Memory unit 84 may include volatile computer-readable media (e.g., dynamic random access memory - DRAM) that stores the encoded data / signals / instructions accessed or generated by processor 82 during operation. Input device 86 may include a computer keyboard, mouse, microphone, etc., and includes corresponding hardware interfaces and / or adapters that allow users to input data and / or instructions into computer system 80.Output device 88 may include a display device (e.g., a monitor and speakers) and a hardware interface / adapter (e.g., a graphics card) to enable the corresponding computer device to communicate data to the user. In some embodiments, input and output devices 86 to 88 share common hardware (e.g., a touchscreen). Storage device 92 includes computer-readable media to enable non-volatile storage, retrieval, and writing of software instructions and / or data. Exemplary storage devices include disk and optical disk and flash memory devices, as well as removable media such as CD and / or DVD optical discs and drives. Network adapter 94 enables computer system 80 to connect to an electronic communication network (e.g., network 15 in FIG. 1) and / or other devices / computer systems.

[0074] Controller hub 90 generally represents multiple system, peripheral, and / or chipset buses, and / or all other circuitry that enables communication between processor 82 and the remaining hardware components of computer system 80. For example, controller hub 90 may include a memory controller, an input / output (I / O) controller, and an interrupt controller. Depending on the hardware manufacturer, some such controllers may be incorporated into a single integrated circuit and / or integrated with processor 82. In another instance, controller hub 90 may include a northbridge connecting processor 82 to memory 84 and / or a southbridge connecting processor 82 to devices 86, 88, 92, and 94.

[0075] The exemplary systems and methods described above enable efficient detection of computer security threats (e.g., malware and intrusions). Embodiments of the present invention address a key technical problem of artificial intelligence (AI), namely that AI systems are typically trained on one set of data and then applied to another set of data. Although AI has strong generalization capabilities, it is not known a priori how well a trained AI system will perform on data not encountered during training. Some embodiments of the present invention address this problem via a feature selection process, in which an initial set of features characterizing input data is reduced to a selected subset of features that may be robust to novelty.

[0076] Some embodiments rely on observations that the real-world data collected for automatic classification and / or anomaly detection contains two distinct contributions or components. One component considered a "signal" in this paper includes specific aspects of the data targeted by the corresponding classifier / detector. In the instance of malware detection, the signal sought may include the behavior of the malware. Other components can be described as "background" and include other aspects of the data or the monitored system that are unrelated to or contain less information about the aforementioned signals. Using an analogy from image processing, the distinction between signal and background can be easily understood. The task of an automatic image segmentation system could be to determine whether an input image displays a bicycle—the signal in the current analogy.However, real images almost never show isolated bicycles; instead, they are surrounded by other objects such as buildings, people, other vehicles, etc., which act as the background. When an image is presented, the segmentation system may extract a set of image features and then feed the corresponding features into a classifier configured to determine whether the corresponding image shows a bicycle. However, it is often unclear which corresponding image features are actually characteristics of a bicycle and which are characteristics of people, buildings, or specific landscapes surrounding the bicycle in the corresponding training images. When the extracted features are too sensitive to the background or too informative, the corresponding classifier may fail when presented with images of types it has not encountered during training (e.g., bicycles on a mountain road). Specification 11 / 13 pages 15 CN 121816573 A

[0077] Similarly, in computer security applications, detectors may be trained to detect threats (e.g., malware) by monitoring software behavior. However, data collected for monitoring purposes encodes both occasional malicious activity (i.e., the sought signal) and the benign normal behavior of the user of the corresponding computer system (hereinafter, the background or environment of the corresponding signal). As in the image processing example, the detector module can extract a set of features from the collected data, but it is not prior to know which features contain more information about malicious activity compared to the user's legitimate behavior. When the extracted features are too sensitive to the user's benign behavior, the corresponding malware detector may fail when the task is to analyze user data that was not seen during training.

[0078] Some embodiments rely on a feature selection process that attempts to pick out features that contain more information about the signal than the background. This process may benefit from the observation of dividing the training data into multiple individual corpora, where each different corpus primarily contains samples with different types of backgrounds, while also containing signals of the same type as the content of other corpora. In the image processing example, one corpus may consist primarily of images of bicycles in urban environments, and another corpus may consist primarily of images of bicycles on rural roads. A similar approach may be applicable to computer security applications because the way a computer is used may vary significantly from user to user. For example, compared to an accountant, an engineer may open other types of applications and access other online resources. Meanwhile, malware (i.e., "signals") may operate in a manner largely independent of the user.

[0079] Based on such observations, some embodiments of the present invention divide the collected training data into different corpora 20a to d, such that each corpus can primarily contain different types of background, i.e., normal or benign user behavior. Examples include placing data collected from different users, different computers, different households, different enterprise customers, different network domains, etc., into separate training corpora.In other exemplary embodiments, data collected at different time intervals (working hours and evenings, weekends, and holidays) are placed into different training corpora. Other exemplary criteria for dividing the collected data into different corpora may include device type (e.g., desktop computers vs. laptop computers vs. mobile computing devices, Microsoft Windows® vs. other operating systems), geographic location (e.g., each country or region placed into different corpora).

[0080] Then, some embodiments implement a feature selection process in which each available feature is ranked according to its frequency distribution across each training corpus. Features whose distribution is relatively stable across multiple corpora may be ranked higher than other features based on the assumption that they may be better at characterizing signal components. A reduced feature set including the higher-ranked features is selected and used to train a classifier for detecting computer security threats.

[0081] Feature selection has been a major focus of the machine learning community. Many feature selection strategies are known in the field. However, conventional feature selection typically involves determining how relevant the selected features are to the output (actual or expected) of the classifier. In other words, some features are better at distinguishing between different categories (e.g., benign / malicious) than others, and are therefore considered more useful for the corresponding classification task. Another conventional feature selection strategy determines the correlation between features across the training corpus. A relatively strong correlation indicates that two features encode roughly the same information, and therefore one of them can be discarded from the feature set. Crucially, all such conventional feature selection processes typically use a single training corpus.

[0082] In contrast, some embodiments use multiple carefully constructed training corpora, and the feature selection process described herein explicitly depends on a specific number of corpora, such as the distribution of the selected features across the selected corpora. Furthermore, corpus construction forms an important part of the feature selection process because the criteria used to partition the collected data samples into different corpora may ultimately determine which features are selected into the reduced feature set. In other words, the contents of the reduced feature set may vary depending on what is considered the “background” of the corresponding training data. Furthermore, feature selection according to some embodiments of the present invention does not concern itself with the output of the classifier and / or the actual class of the corresponding data sample, but rather relies on the eigenvalues ​​themselves.

[0083] Some important benefits arising from feature selection as described herein include a significant reduction in computational costs associated with both training and using a smaller feature set. In machine learning, it is known that computational costs are approximately proportional to the square of the size of the feature vector, so reducing the number of features by only 25% can potentially reduce the associated computational burden by nearly half.Such significant reductions have a positive impact on user experience and facilitate the implementation of reliable computer security software on devices lacking the processor and memory resources of desktop or laptop computers (e.g., smartphones and Internet of Things (IoT) devices). Further potential benefits include shorter time-to-market and facilitate frequent updates of security software to keep up with rapidly evolving threats. Significantly, the reduction in computational burden is further accompanied by improved detection performance, as demonstrated by computer experiments.

[0084] Another advantage of some of the embodiments described herein is the significant robustness of the detector when presented with different, previously unseen data. Because the feature selection process presented herein deliberately favors features containing less background / environmental information, the resulting system performs well on novel tasks, as demonstrated in computer experiments.

[0085] Although the above discussion focuses primarily on computer security applications, the described systems and methods can be adapted to other applications, such as image processing (e.g., segmentation, search, annotation, deepfake detection) and natural language processing (e.g., online fraud detection, automatically generated text and chatbots, author attribute detection, etc.). For example, some embodiments may divide the corpus of chatbot-generated text into different training corpora based on the type / identity of the chatbot that generated the corresponding text (e.g., ChatGPT® vs. LlamaChat®), where each different corpus primarily consists of samples from a single source / type of chatbot. The chatbot detector can then be trained to use a reduced feature set selected according to some of the methods described herein. Such detectors are expected to perform well when presented with automatically generated text from unknown sources.

[0086] Furthermore, some embodiments may help extend machine learning / AI methods currently used in fields such as natural language processing to other areas of information technology. Examples include automated analysis, annotation, reverse engineering, and computer code generation, which are increasingly being introduced from natural language processing methods. One exemplary embodiment of the invention may construct multiple training corpora, where each different corpus primarily comprises code written in different programming languages ​​(e.g., C++ vs. Java vs. Python), different programming paradigms (e.g., imperative vs. declarative vs. object-oriented), different tagging standards, etc. Applying the feature selection process described herein can result in robust code embeddings that are relatively insensitive to language and are therefore advantageous in applications such as malware detection and classification.

[0087] Those skilled in the art will recognize that the above embodiments can be modified in various ways without departing from the scope of the invention. Therefore, the scope of the invention should be determined by the appended claims and their legal equivalents.Instruction manual, page 13 / 13, 17 CN 121816573 A, Figure 1; Instruction manual, Figure 1 / 7, page 18 CN 121816573 A, Figure 2; Instruction manual, Figure 2 / 7, page 19 CN 121816573 A, Figure 3; Instruction manual, Figure 3 / 7, page 20 CN 121816573 A, Figure 4; Instruction manual, Figure 4 / 7, page 21 CN 121816573 A, Figure 6; Instruction manual, Figure 5 / 7, page 22 CN 121816573 A, Figure 7; Instruction manual, Figure 6 / 7, page 23 CN 121816573 A, Figure 8; Instruction manual, Figure 7 / 7, page 24 CN 121816573 A.

Claims

1. A computer system comprising at least one hardware processor configured to: Selecting a reduced subset of features from a plurality of features that can be used to characterize the data sample, wherein selecting the reduced subset of features includes: The data samples obtained from multiple computing devices are divided into multiple training corpora; Candidate features are selected from the plurality of features. Determine the first frequency distribution of the feature values ​​of the candidate features across members of a first training corpus in the plurality of training corpora. Determine the second frequency distribution of the feature values ​​of the candidate features on the second training corpus among the multiple training corpora, and Based on the similarity between the first frequency distribution and the second frequency distribution, determine whether to include the candidate feature in the reduced feature subset; and In response to selecting the reduced feature subset, a threat detector is trained to determine whether a target data sample indicates a computer security threat based on the reduced feature subset.

2. The computer system of claim 1, wherein the at least one hardware processor is configured to divide the set of data samples into the plurality of training corpora according to the location of the computing device providing each corresponding data sample.

3. The computer system of claim 1, wherein the at least one hardware processor is configured to divide the set of data samples into the plurality of training corpora according to the identity of the user of the computing device providing each corresponding data sample.

4. The computer system of claim 1, wherein the at least one hardware processor is configured to divide the set of data samples into the plurality of training corpora according to the acquisition time of each corresponding data sample.

5. The computer system of claim 1, wherein the at least one hardware processor is configured to divide the set of data samples into the plurality of training corpora according to the device type of the computing device providing each corresponding data sample.

6. The computer system of claim 1, wherein the plurality of computing devices are divided among a plurality of company owners, and wherein the at least one hardware processor is configured to divide the set of data samples into the plurality of training corpora according to the owner of the computing device providing each corresponding data sample.

7. The computer system of claim 1, wherein the at least one hardware processor is configured to partition the set of data samples into the plurality of training corpora according to a software configuration file of the computing device providing each corresponding data sample, the software configuration file comprising a set of computer programs installed for execution on the corresponding computing device.

8. The computer system of claim 1, wherein determining whether to include the candidate feature in the reduced feature subset further comprises: A plurality of similarity measures are determined, each of the plurality of similarity measures quantifying the similarity between a pair of frequency distributions of the values ​​of the candidate feature, and each of the pair of probability measures is evaluated on different corpora in the plurality of training corpora; and The candidate features are selected into the reduced feature subset based on the average of the multiple similarity measurements.

9. The computer system of claim 8, wherein the at least one hardware processor is configured to further select the candidate features into the reduced feature subset based on the degree of dispersion of the plurality of similarity measurements.

10. The computer system of claim 1, wherein determining whether to include the candidate feature in the reduced feature subset further comprises: For each of the plurality of features, evaluate the feature-specific frequency distribution of the feature value of the corresponding feature on the first training corpus; The plurality of features are ranked according to the evaluated feature-specific frequency distribution; and Based on the ranking results, the candidate features are selected into the reduced feature subset.

11. The computer system of claim 1, wherein the plurality of features are automatically constructed by a machine learning process, the machine learning process including training another threat detector to identify members of a set of data samples indicating a computer security threat.

12. A computer security method comprising employing at least one hardware processor of a computer system to: Selecting a reduced subset of features from a plurality of features that can be used to characterize the data sample, wherein selecting the reduced subset of features includes: The data samples obtained from multiple computing devices are divided into multiple training corpora; Candidate features are selected from the plurality of features. Determine the first frequency distribution of the feature values ​​of the candidate features across members of a first training corpus in the plurality of training corpora. Determine the second frequency distribution of the feature values ​​of the candidate features on the second training corpus among the multiple training corpora, and Based on the similarity between the first frequency distribution and the second frequency distribution, determine whether to include the candidate feature in the reduced feature subset; and In response to selecting the reduced feature subset, a threat detector is trained to determine whether a target data sample indicates a computer security threat based on the reduced feature subset.

13. The method of claim 12, further comprising dividing the set of data samples into the plurality of training corpora based on the location of the computing device providing each corresponding data sample.

14. The method of claim 12, further comprising dividing the set of data samples into the plurality of training corpora based on the identity of the user of the computing device providing each corresponding data sample.

15. The method of claim 12, further comprising dividing the set of data samples into the plurality of training corpora according to the acquisition time of each corresponding data sample.

16. The method of claim 12, further comprising dividing the set of data samples into the plurality of training corpora according to the device type of the computing device providing each corresponding data sample.

17. The method of claim 12, further comprising partitioning the set of data samples into the plurality of training corpora according to a software configuration file of the computing device providing each corresponding data sample, the software configuration file including a set of computer programs installed for execution on the corresponding computing device.

18. The method of claim 12, wherein the plurality of computing devices are divided among a plurality of company owners, the method comprising dividing the set of data samples into the plurality of training corpora according to the owner of the computing device providing each corresponding data sample.

19. The method of claim 12, wherein determining whether to include the candidate feature in the reduced subset of features further comprises: A plurality of similarity measures are determined, each of the plurality of similarity measures quantifying the similarity between a pair of frequency distributions of the values ​​of the candidate feature, and each of the pair of probability measures is evaluated on different corpora in the plurality of training corpora; and The candidate features are selected into the reduced feature subset based on the average of the multiple similarity measurements.

20. The method of claim 19, further comprising selecting the candidate features into the reduced feature subset based on the degree of dispersion of the plurality of similarity measurements.

21. The method of claim 12, wherein determining whether to include the candidate feature in the reduced feature subset further comprises: For each of the plurality of features, evaluate the feature-specific frequency distribution of the feature value of the corresponding feature on the first training corpus; The plurality of features are ranked according to the evaluated feature-specific frequency distribution; and Based on the ranking results, the candidate features are selected into the reduced feature subset.

22. The method of claim 12, wherein the plurality of features are automatically constructed by a machine learning process, the machine learning process including training another threat detector to identify members of a set of data samples indicating the computer security threat.

23. A non-transitory computer-readable medium storing instructions that, when executed by at least one hardware processor of a computer system, cause the computer system to: Selecting a reduced subset of features from a plurality of features that can be used to characterize the data sample, wherein selecting the reduced subset of features includes: The data samples obtained from multiple computing devices are divided into multiple training corpora; Candidate features are selected from the plurality of features. Determine the first frequency distribution of the feature values ​​of the candidate features across members of a first training corpus in the plurality of training corpora. Determine the second frequency distribution of the feature values ​​of the candidate features on the second training corpus among the multiple training corpora, and Based on the similarity between the first frequency distribution and the second frequency distribution, determine whether to include the candidate feature in the reduced feature subset; and In response to selecting the reduced feature subset, a threat detector is trained to determine whether a target data sample indicates a computer security threat based on the reduced feature subset.