Malicious software open set identification method and device

By using multimodal feature extraction and fusion learning, combined with radial basis function support vector classifier and multi-stream threshold network, the problem of low accuracy and insufficient identification of unknown samples in existing technologies for malware detection is solved, and efficient identification and defense against malware is achieved.

CN120850285APending Publication Date: 2025-10-28GUANGDONG POLYTECHNIC NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510890103.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing malware detection methods cannot effectively identify unknown family samples not seen during the training phase, and relying on a single feature makes it difficult to fully characterize the complex behavior of malware, resulting in high false positive rates and low detection accuracy.

Method used

Multimodal feature extraction and fusion learning are employed, combining radial basis function support vector classifier and multi-stream threshold network. Decision boundaries are constructed by aligning and weighting features through triplet loss function, and static flow, dynamic flow and API flow classifiers are used for identification.

Benefits of technology

It significantly improves the accuracy and generalization ability of malware identification, effectively identifies unknown categories, and enhances detection performance and security defense capabilities in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120850285A_ABST
    Figure CN120850285A_ABST
Patent Text Reader

Abstract

The invention discloses a malicious software open set identification method and device, and the method comprises the steps: obtaining to-be-detected sample data, carrying out the feature extraction of the to-be-detected sample data, and obtaining the multi-modal features of each to-be-detected sample in the to-be-detected sample data; performing alignment processing and weighted fusion processing on the multi-modal features of the to-be-detected samples based on a preset triple loss function to obtain fusion feature vectors of the to-be-detected samples; constructing a radial basis function support vector classifier, identifying the fusion feature vector based on the radial basis function support vector classifier, and determining a decision boundary; and constructing a multi-stream threshold network based on a preset static stream classifier, a preset dynamic classifier and a preset AP I stream classifier, identifying the multi-modal features based on the multi-stream threshold network, obtaining a category label of each to-be-detected sample, and generating a detection result of the to-be-detected sample data based on the decision boundary and the category labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of malware detection technology, and in particular to a method and apparatus for identifying open sets of malware. Background Technology

[0002] Currently, malware (full name "malicious software program") refers to software specifically designed and written to steal, damage, or illegally control computer systems, networks, or user data. Its core characteristic is that it secretly performs harmful operations on the victim's device without the user's consent or knowledge.

[0003] Existing malware detection methods suffer from several drawbacks. First, they are based on the closed-world assumption, which fails to effectively identify unknown family samples not seen during the training phase, leading to high false positive rates. Second, existing detection methods often rely on single features (such as static code or dynamic behavior), making it difficult to comprehensively characterize the complex behavior of malware. Furthermore, fixed threshold strategies cannot adapt to the dynamic distribution of high-dimensional feature spaces, resulting in blurred classification boundaries and consequently low malware detection accuracy. Summary of the Invention

[0004] This invention provides a method and apparatus for identifying open sets of malicious software, thereby improving the accuracy of malicious software detection.

[0005] To address the aforementioned technical problems, this invention provides a method for identifying open sets of malicious software, comprising:

[0006] Acquire the sample data to be detected, extract features from the sample data to be detected, and obtain the multimodal features of each sample to be detected in the sample data to be detected;

[0007] Based on the preset triplet loss function, the multimodal features of each sample to be detected are aligned and weighted and fused to obtain the fused feature vector of each sample to be detected.

[0008] A radial basis function support vector classifier is constructed, and the fused feature vector is identified based on the radial basis function support vector classifier to determine the decision boundary;

[0009] A multi-stream threshold network is constructed based on a preset static stream classifier, a preset dynamic stream classifier, and a preset API stream classifier. The multi-modal features are identified based on the multi-stream threshold network to obtain the category label of each sample to be detected. The detection result of the sample data to be detected is generated based on the decision boundary and the category label.

[0010] This invention effectively improves the accuracy and generalization ability of identification in environments with unknown malicious samples. By extracting and fusing multimodal features such as static, dynamic, and API call sequences, the model gains a more comprehensive perception of malicious software samples. The introduction of a triplet loss function enhances feature alignment and discriminative power, improving inter-class separability. The use of radial basis function support vector machines (SVC) to construct decision boundaries, combined with multi-stream thresholding networks to fuse classification results, achieves accurate identification of known categories and effective rejection of unknown categories, improving identification accuracy and significantly enhancing detection performance and security defense capabilities in open-set environments.

[0011] Furthermore, the sample data to be tested includes executable files and sandbox logs of several samples to be tested; the process of acquiring the sample data to be tested, performing feature extraction on the sample data to be tested, and obtaining the multimodal features of each sample to be tested in the sample data to be tested includes:

[0012] Obtain the executable file of the sample to be detected, extract features from the executable file, and obtain the static features of the sample to be detected;

[0013] Feature extraction is performed on the sandbox logs to obtain the running behavior data of the sample to be detected, and dynamic features of the sample to be detected are generated based on the running behavior data;

[0014] Feature extraction is performed on the sandbox logs to obtain the API call sequence features of the sample to be detected;

[0015] The static features, dynamic features, and API call sequence features constitute the multimodal features of the sample to be detected.

[0016] This invention clarifies the composition of multimodal features by limiting the source of sample data and its corresponding feature extraction path, including static features corresponding to executable files, dynamic behaviors extracted from sandbox logs, and API call sequence features. This feature extraction method not only covers information on malware in both static structure and runtime behavior, but also further introduces temporal behavioral features from the API call level. This results in multimodal features with higher information density and discriminative power, providing a more reliable input foundation for subsequent discrimination modeling, thereby improving the overall recognition accuracy and robustness of the system.

[0017] Furthermore, the multimodal features of each sample to be detected are aligned and weighted and fused based on a preset triplet loss function to obtain the fused feature vector of each sample to be detected, including:

[0018] The static features, dynamic features, and API call sequences of each sample to be detected are normalized to obtain normalized multimodal features.

[0019] The normalized multimodal features are mapped based on a preset metric learning network to generate embedding vectors for each sample to be detected.

[0020] The embedding vector is optimized based on a preset triplet loss function to generate a fused feature vector for each sample to be detected.

[0021] This invention addresses the issues of inconsistent scales and semantic misalignment in multi-source heterogeneous features by normalizing multi-modal features, implementing metric learning mapping, and optimizing the triplet loss function. This significantly improves the consistency of fused feature representation and discriminative performance. The triplet loss function effectively minimizes the distance between similar samples and maximizes the distance between dissimilar samples, ensuring clear boundaries between different categories in the embedding space and avoiding ambiguous and overlapping clustering results. This enhances the model's classification and discriminative capabilities when dealing with boundary samples and approximate samples.

[0022] Furthermore, the construction of a radial basis function support vector classifier, and the identification of the fused feature vector based on the radial basis function support vector classifier to determine the decision boundary, includes:

[0023] Constructing a dynamic classifier based on radial basis function support vector classifier;

[0024] The fused feature vector is input into a pre-trained radial basis function support vector classifier to obtain the classification confidence and classification label of each sample to be detected;

[0025] Calculate the geometric distance from the sample to be detected to the classification label based on the classification label;

[0026] The decision boundary is determined based on the classification confidence level and geometric distance.

[0027] This invention introduces an RBF kernel support vector classifier and constructs a dynamic decision-making mechanism based on classification confidence and the geometric distance from the sample to the decision boundary, enabling adjustable control over the discrimination boundary of open sets of malware. During the model prediction phase, this mechanism can determine whether to reject a sample based on its classification confidence and relative position, effectively improving the ability to identify unknown categories, alleviating the problem of "overfitting known classes" in traditional closed-set classification methods, and enhancing the system's flexibility and anti-interference capabilities in dealing with new or variant malware in real-world deployment environments.

[0028] Furthermore, the multi-stream thresholding network is constructed based on a preset static stream classifier, a preset dynamic classifier, and a preset API stream classifier. The multi-stream thresholding network is used to identify the multimodal features, obtain the category labels for each sample to be detected, and generate the detection results for the sample data based on the decision boundary and the category labels. This includes:

[0029] Based on the preset static flow classifier, preset dynamic classifier and preset API flow classifier, static features, dynamic features and API call sequence features are identified respectively, and static identification results, dynamic identification results and API identification results are obtained;

[0030] The static recognition results, dynamic recognition results, and API recognition results are weighted and fused based on a preset soft voting mechanism to obtain the category label of each sample to be detected, and the detection result of the sample data to be detected is generated based on the decision boundary and the category label.

[0031] This invention constructs three sub-classifiers—static, dynamic, and API stream—and introduces a soft voting mechanism to fuse the results from the three streams, effectively improving the stability and robustness of the classification process. This multi-stream fusion strategy leverages the recognition advantages of each modality in specific situations, achieving multi-faceted information complementarity and significantly reducing the misclassification rate caused by missing or distorted sample features in a single classifier. Furthermore, combined with the dynamic decision boundary constructed in the superior claims, it further enhances the ability to handle edge samples and improves the accuracy of malware identification systems in complex environments.

[0032] Secondly, the present invention provides a malware open set identification device, comprising: a feature extraction module, a feature fusion module, a decision adjustment module, and an identification module;

[0033] The feature extraction module is used to acquire the sample data to be detected, perform feature extraction on the sample data to be detected, and acquire the multimodal features of each sample to be detected in the sample data to be detected.

[0034] The feature fusion module is used to perform alignment and weighted fusion processing on the multimodal features of each sample to be detected based on a preset triplet loss function, so as to obtain the fused feature vector of each sample to be detected.

[0035] The decision adjustment module is used to construct a radial basis function support vector classifier, and to identify the fused feature vector based on the radial basis function support vector classifier to determine the decision boundary;

[0036] The recognition module is used to construct a multi-stream threshold network based on a preset static stream classifier, a preset dynamic classifier, and a preset API stream classifier, recognize the multimodal features based on the multi-stream threshold network, obtain the category label of each sample to be detected, and generate the detection result of the sample data to be detected based on the decision boundary and the category label.

[0037] Furthermore, the feature extraction module is used to obtain the sample data to be detected, which includes executable files and sandbox logs of several samples to be detected; the step of obtaining the sample data to be detected and performing feature extraction on the sample data to be detected to obtain the multimodal features of each sample to be detected in the sample data to be detected includes:

[0038] Obtain the executable file of the sample to be detected, extract features from the executable file, and obtain the static features of the sample to be detected;

[0039] Feature extraction is performed on the sandbox logs to obtain the running behavior data of the sample to be detected, and dynamic features of the sample to be detected are generated based on the running behavior data;

[0040] Feature extraction is performed on the sandbox logs to obtain the API call sequence features of the sample to be detected;

[0041] The static features, dynamic features, and API call sequence features constitute the multimodal features of the sample to be detected.

[0042] Furthermore, the feature fusion module is used to perform alignment and weighted fusion processing on the multimodal features of each sample to be detected based on a preset triplet loss function, to obtain the fused feature vector of each sample to be detected, including:

[0043] The static features, dynamic features, and API call sequences of each sample to be detected are normalized to obtain normalized multimodal features.

[0044] The normalized multimodal features are mapped based on a preset metric learning network to generate embedding vectors for each sample to be detected.

[0045] The embedding vector is optimized based on a preset triplet loss function to generate a fused feature vector for each sample to be detected.

[0046] Furthermore, the decision adjustment module is used to construct a radial basis function support vector classifier, and to identify the fused feature vector based on the radial basis function support vector classifier to determine the decision boundary, including:

[0047] Constructing a dynamic classifier based on radial basis function support vector classifier;

[0048] The fused feature vector is input into a pre-trained radial basis function support vector classifier to obtain the classification confidence and classification label of each sample to be detected;

[0049] Calculate the geometric distance from the sample to be detected to the classification label based on the classification label;

[0050] The decision boundary is determined based on the classification confidence level and geometric distance.

[0051] Furthermore, the recognition module is used to construct a multi-stream threshold network based on a preset static stream classifier, a preset dynamic classifier, and a preset API stream classifier; to recognize the multimodal features based on the multi-stream threshold network; to obtain the category label of each sample to be detected; and to generate the detection result of the sample data to be detected based on the decision boundary and the category label, including:

[0052] Based on the preset static flow classifier, preset dynamic classifier and preset API flow classifier, static features, dynamic features and API call sequence features are identified respectively, and static identification results, dynamic identification results and API identification results are obtained;

[0053] The static recognition results, dynamic recognition results, and API recognition results are weighted and fused based on a preset soft voting mechanism to obtain the category label of each sample to be detected, and the detection result of the sample data to be detected is generated based on the decision boundary and the category label. Attached Figure Description

[0054] Figure 1 A flowchart illustrating a method for identifying open sets of malicious software provided in an embodiment of the present invention;

[0055] Figure 2 This is a distribution diagram of the number of samples in the BODMAS dataset provided in an embodiment of the present invention;

[0056] Figure 3 This is a schematic diagram illustrating the known / unknown division of the first 80 families in the BODMAS dataset provided in this embodiment of the invention;

[0057] Figure 4 A flowchart for extracting dynamic features and API function sequences provided in embodiments of the present invention;

[0058] Figure 5 A two-dimensional visualization diagram illustrating the metric learning model process of the feature space provided in this embodiment of the invention;

[0059] Figure 6 A schematic diagram of the EMST framework provided in an embodiment of the present invention;

[0060] Figure 7 This is a schematic diagram of the feature space for t-SNE two-dimensional visualization provided in an embodiment of the present invention;

[0061] Figure 8 This is a schematic diagram of the prediction confusion matrix of each classifier provided in the embodiments of the present invention;

[0062] Figure 9This is a schematic diagram of the structure of a malware open set identification device provided in an embodiment of the present invention. Detailed Implementation

[0063] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0064] The terms "first" and "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.

[0065] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0066] Example 1

[0067] See Figure 1 , Figure 1 This is a flowchart illustrating a method for identifying open sets of malware according to an embodiment of the present invention. The embodiment of the present invention provides a method for identifying open sets of malware, including steps 101 to 104, as detailed below:

[0068] Step 101: Obtain the sample data to be detected, perform feature extraction on the sample data to be detected, and obtain the multimodal features of each sample to be detected in the sample data to be detected;

[0069] In this embodiment, the sample data to be detected includes executable files and sandbox logs of several samples to be detected; the step of acquiring the sample data to be detected, performing feature extraction on the sample data to be detected, and obtaining the multimodal features of each sample to be detected in the sample data to be detected includes:

[0070] Obtain the executable file of the sample to be detected, extract features from the executable file, and obtain the static features of the sample to be detected;

[0071] Feature extraction is performed on the sandbox logs to obtain the running behavior data of the sample to be detected, and dynamic features of the sample to be detected are generated based on the running behavior data;

[0072] Feature extraction is performed on the sandbox logs to obtain the API call sequence features of the sample to be detected;

[0073] The static features, dynamic features, and API call sequence features constitute the multimodal features of the sample to be detected.

[0074] Please refer to Figure 2 , Figure 2 This is a distribution chart of the number of samples in the BODMAS dataset provided in this embodiment of the invention. The horizontal axis represents the range of sample numbers, the vertical axis represents the total number of samples, and the number at the top of the bar represents the number of families.

[0075] Please refer to Figure 3 , Figure 3 This diagram illustrates the known / unknown classification of the first 80 families in the BODMAS dataset provided in this embodiment of the invention. The families are sorted from most to least represented by the number of samples; blue indicates known family labels, and red indicates unknown family labels.

[0076] In this embodiment, the executable files (.exe files) to be detected are obtained through user clients, network monitoring devices, or sandbox dynamic analysis systems. Specifically, the BODMAS dataset is used, which contains 134,435 Windows executable files, covering 14 types of malware (such as Trojans, worms, and ransomware) and 581 families. Its long-tail distribution (7 families with more than 2,000 samples and 496 families with fewer than 50 samples) simulates the real malware ecosystem. Sorted by the number of family samples, the top 80% of families (464) are classified as known categories (…). Figure 3 The blue area), the latter 20% of families (117) are of unknown category ( Figure 3 (Red area). The training set (60%), validation set (20%), and test set (20%) are divided chronologically to simulate the temporal evolution characteristics of malware.

[0077] In this embodiment, the executable file of the sample to be tested is first obtained, and static analysis is performed using the Ember tool. The extracted static features include: (1) byte histogram features, which form a 256-dimensional feature vector by statistically analyzing the frequency of 256 byte values ​​in the file, used to characterize the binary distribution characteristics of the file; (2) imported function features, which record the library function information called by the PE file (such as the CreateProcess function in kernel32.dll), reflecting the functional intent of the program; (3) section features, which extract the name, size and entropy value of each section (such as .text, .data), where the entropy value is calculated using the Shannon entropy formula to quantify the randomness of the section data. Finally, the above features are standardized into a 2,381-dimensional NumPy array format.

[0078] In this embodiment, runtime behavior data is extracted by parsing the VirusTotal sandbox logs: (1) file operation features, recording the path and timestamp of files created / modified by the samples (e.g., C:\temp\1.exe); (2) process behavior features, constructing a process tree and capturing abnormal termination events; (3) registry operation features, extracting sensitive key value modification records (e.g., write operations to HKLM\Software\Microsoft). The TF-IDF algorithm is used to weight the behavior strings, and after filtering high-frequency behavior patterns, a normalized sparse vector is generated.

[0079] In this embodiment, the TianQiong sandbox XML logs are parsed to extract the temporal API call chain (such as the CreateFileW→WriteFile→RegSetValueEx sequence). API word vectors are trained using a CBOW model: the sliding window size is set to 3, the context API calls are taken as input, and a 128-dimensional semantic vector is output. Through neural network training, APIs with similar contexts (such as CreateFileW and WriteFile) are made close in distance in the vector space, thereby capturing the temporal semantic relationships between APIs.

[0080] In this embodiment, the static features (2,381 dimensions), dynamic features (TF-IDF sparse vectors), and API sequence features (128-dimensional word vectors) are ultimately fused in a multimodal manner to form a joint feature representation of the sample. This method significantly improves the accuracy of malicious code detection and the ability to identify adversarial examples by combining three features: static file attributes, runtime behavior patterns, and API call semantics.

[0081] In this embodiment, by defining the source of sample data and its corresponding feature extraction path, the composition method of multimodal features is clarified, including static features corresponding to executable files, dynamic behaviors extracted from sandbox logs, and API call sequence features. This feature extraction method not only covers information on malware in both static structure and runtime behavior dimensions, but also further introduces temporal behavioral features from the API call level, resulting in multimodal features with higher information density and discriminative power. This provides a more reliable input basis for subsequent discrimination modeling, thereby improving the overall recognition accuracy and robustness of the system.

[0082] Step 102: Based on the preset triplet loss function, perform alignment and weighted fusion processing on the multimodal features of each sample to be detected to obtain the fused feature vector of each sample to be detected.

[0083] In this embodiment, the multimodal features of each sample to be detected are aligned and weighted based on a preset triplet loss function to obtain the fused feature vector of each sample to be detected, including:

[0084] The static features, dynamic features, and API call sequences of each sample to be detected are normalized to obtain normalized multimodal features.

[0085] The normalized multimodal features are mapped based on a preset metric learning network to generate embedding vectors for each sample to be detected.

[0086] The embedding vector is optimized based on a preset triplet loss function to generate a fused feature vector for each sample to be detected.

[0087] Please refer to Figure 4 , Figure 4 A flowchart for extracting dynamic features and API function sequences provided in embodiments of the present invention.

[0088] In this embodiment, firstly, the multimodal features of each sample to be detected are normalized, specifically including static features, dynamic behavioral features, and API call sequence features. Static features are extracted using Ember to obtain a 2381-dimensional feature vector, which, after normalization, is directly used as static input. Dynamic features are vectorized using TF-IDF and then L2 normalized to obtain a 1000-dimensional dense vector. The API call sequence represents each API function as a 128-dimensional vector, with the sequence length truncated or padded to 100, forming a 128×100 matrix. This matrix is ​​input to a Bidirectional Long Short-Term Memory (BiLSTM) network, and its temporal features are extracted and then pooled to output a 128-dimensional temporal feature vector.

[0089] In this embodiment, the three features are mapped to a unified 256-dimensional shared space through fully connected layers. Static features are mapped through a 2,381→256-dimensional fully connected layer, dynamic features through a 1,000→256-dimensional fully connected layer, and API features through a 128→256-dimensional fully connected layer, thus completing dimensional alignment and semantic fusion.

[0090] Please refer to Figure 5 , Figure 5 This is a two-dimensional visualization diagram of the feature space metric learning model process provided in an embodiment of the present invention.

[0091] In this embodiment, the 256-dimensional features of the three modalities mentioned above are input into a metric learning network. A pre-defined triplet loss function is introduced into this network, and combinations of Anchor (anchor sample), Positive (same class sample), and Negative (different class sample) are designed to constrain the embedding space structure learned by the model. By minimizing the Euclidean distance between Anchor and Positive while maximizing the distance between Anchor and Negative, effective modeling of semantic relationships between features is achieved, thereby improving inter-class discriminativeness.

[0092] Finally, the embedding vector optimized by the triplet loss function becomes the fused feature vector, possessing good clustering properties and open set separability. It can be used as input to subsequent classifiers to improve the detection capability for unknown categories. This process not only achieves a unified representation of multimodal features but also enhances the model's ability to recognize fuzzy boundary samples through metric learning, effectively improving the generalization performance and robustness of the overall detection system.

[0093] In this embodiment, a weighted triplet loss function based on a hard sample mining strategy is adopted. The specific process is as follows: First, the multimodal fusion feature vector of each sample to be detected is used as input and fed into a metric learning network consisting of three fully connected layers. The network output is the mapping result of the feature embedding function φ(x) to the input vector. During model training, anchor-positive-negative sample triplets (anchor, positive, negative sample) are randomly sampled in batches, and the weight of each triplet is calculated.

[0094]

[0095] Here, ε is a manually set scaling factor used to amplify the impact of small sample difficulty differences on the weights; φ(·) is the embedding mapping of the metric learning network; and ||·||² represents the Euclidean distance. This weight assigns higher priority to "difficult negative samples" (negative samples that are closer to the anchor point) and lower priority to "easy negative samples," thereby achieving dynamic hard sample mining.

[0096] Next, the model uses the weighted triplet loss function as the optimization objective and summarizes the loss for all triplet samples. The loss is defined as follows:

[0097]

[0098] Here, γ is a user-defined margin parameter used to ensure that the distance between negative samples is at least γ greater than the distance between positive samples.

[0099] In this embodiment, by iteratively minimizing the weighted loss function, the network can selectively bring similar samples closer together and distance easily confused dissimilar samples from the embedding space, achieving focused optimization of boundary and difficult-to-distinguish samples. Ultimately, the trained metric learning network generates fused feature vectors with better class separation and open set detection capabilities, providing high-quality input for subsequent support vector classifiers and multi-stream thresholding networks. This significantly improves the accuracy and robustness of the method in complex malware open set scenarios.

[0100] In this embodiment, by normalizing multimodal features, performing metric learning mapping, and optimizing the triplet loss, the problems of inconsistent scales and semantic misalignment of multi-source heterogeneous features are solved, significantly improving the consistency of expression and discriminative performance of fused features. The triplet loss function can effectively minimize the distance between samples of the same class and maximize the distance between samples of different classes, ensuring clear boundaries between different categories in the embedding space and avoiding fuzzy and overlapping clustering results, thereby enhancing the model's classification and discriminative ability when facing boundary samples and approximate samples.

[0101] Step 103: Construct a radial basis function support vector classifier, and identify the fused feature vector based on the radial basis function support vector classifier to determine the decision boundary;

[0102] In this embodiment, the step of constructing a radial basis function support vector classifier, and identifying the fused feature vector based on the radial basis function support vector classifier to determine the decision boundary, includes:

[0103] Constructing a dynamic classifier based on radial basis function support vector classifier;

[0104] The fused feature vector is input into a pre-trained radial basis function support vector classifier to obtain the classification confidence and classification label of each sample to be detected;

[0105] Calculate the geometric distance from the sample to be detected to the classification label based on the classification label;

[0106] The decision boundary is determined based on the classification confidence level and geometric distance.

[0107] In this embodiment, firstly, multiple one-to-many RBF kernel support vector classifiers (SVCs) are trained using fused feature vectors and known family labels, where the k-th classifier corresponds to the k-th known family. The classification decision function is expressed as:

[0108]

[0109] Where φ(x) is the feature vector of sample x after Gaussian kernel mapping, w k Let b be the classifier weight vector. k This is a bias term.

[0110] During the testing phase, the fused feature vector of each sample to be detected is input into the pre-trained classifiers to obtain the classification confidence of the sample for each known family label and the corresponding classification label (i.e., the family index corresponding to the highest confidence). Then, based on the classifier corresponding to the classification label, the geometric distance of the sample to the hyperplane is calculated:

[0111]

[0112] In this embodiment, in order to balance high-precision discrimination of known categories with effective elimination of unknown categories, a dual dynamic threshold adjustment mechanism is introduced:

[0113] Using the classification confidence threshold τ k To determine the confidence level Pk of a known class for a sample of class k, a lower limit τ is set. k Only if Pk≥τ k Only when the sample is found to belong to the k-th known family is it considered credibly true; otherwise, it will not be classified into any known class for the time being.

[0114] Using the open set boundary threshold δ k To determine the unknown class, statistically analyze the dk(x) distribution of each classifier's output on the validation set, letting... (μk and σk are the mean and standard deviation of the distribution, respectively). Only when dk(x)≥δk is the sample considered sufficiently far from the decision boundary of the k-th family; otherwise, it is considered a potentially unknown sample.

[0115] In actual judgment, a sample is confirmed as the k-th known family only if it simultaneously satisfies the two conditions Pk≥τk and dk(x)≥δk under the same classifier; if either condition is not met, the sample is marked as an "unknown family" and enters the subsequent multi-stream fusion or manual review process.

[0116] In this embodiment, by introducing an RBF kernel support vector classifier and constructing a dynamic decision-making mechanism based on classification confidence and the geometric distance from the sample to the decision boundary, adjustable control of the discrimination boundary of open sets of malware is achieved. This mechanism can determine whether to reject a sample based on its classification confidence and relative position during the model prediction stage, effectively improving the ability to identify unknown categories, alleviating the problem of "overfitting known classes" in traditional closed-set classification methods, and enhancing the system's flexibility and anti-interference capabilities in dealing with new or variant malware in actual deployment environments.

[0117] Step 104: Construct a multi-stream threshold network based on a preset static stream classifier, a preset dynamic classifier, and a preset API stream classifier. Based on the multi-stream threshold network, identify the multimodal features, obtain the category labels of each sample to be detected, and generate the detection results of the sample data to be detected based on the decision boundary and the category labels.

[0118] In this embodiment, a multi-stream thresholding network is constructed based on a preset static stream classifier, a preset dynamic classifier, and a preset API stream classifier. The multi-stream thresholding network is used to identify the multimodal features, obtain the category labels for each sample to be detected, and generate detection results for the sample data based on the decision boundary and the category labels. This includes:

[0119] Based on the preset static flow classifier, preset dynamic classifier and preset API flow classifier, static features, dynamic features and API call sequence features are identified respectively, and static identification results, dynamic identification results and API identification results are obtained;

[0120] The static recognition results, dynamic recognition results, and API recognition results are weighted and fused based on a preset soft voting mechanism to obtain the category label of each sample to be detected, and the detection result of the sample data to be detected is generated based on the decision boundary and the category label.

[0121] Please refer to Figure 6 , Figure 6 This is a schematic diagram of the EMST framework provided in an embodiment of the present invention.

[0122] In this embodiment, the multi-stream thresholding network specifically includes three parallel sub-streams: a static stream classifier, a dynamic stream classifier, and an API stream classifier. The final identification is achieved through a soft voting mechanism and dual threshold determination. First, the static stream classifier uses a fully connected neural network structure, taking a 256-dimensional fused feature vector as input. After two levels of dimensionality reduction (256→128→64), it outputs the probability distribution of each known family. The dynamic stream classifier, based on a radial basis function kernel support vector machine, calculates decision values ​​for 1,000-dimensional dynamic features and generates probability distributions through Platt scaling. The API stream classifier processes 128×100 API sequence features through a BiLSTM network (hidden layer size = 64) and outputs class probabilities through a Softmax layer. After each of the three sub-streams independently completes its prediction, a soft voting mechanism is used to weight their output probabilities according to w. m A weighted summation is performed, where the weights of each stream are dynamically allocated based on their respective F1 scores in the validation set to fully utilize the discriminative advantages of different modalities. Subsequently, the overall confidence level P after weighted fusion is calculated. ensemble (c * |x)≥τ c *A dual determination is made using the confidence threshold τc and the boundary distance threshold δc for this category: Only when the overall confidence of sample x is not less than τc and its normalized boundary distance dc is not greater than δc is it determined to be a known family c; otherwise, it is marked as "unknown family malware", triggering a real-time alarm and recording detailed behavior logs for subsequent security analysis and response.

[0123] In this embodiment, by constructing three sub-classifiers—static, dynamic, and API stream—and introducing a soft voting mechanism to fuse the results from the three streams, the stability and robustness of the classification process are effectively improved. This multi-stream fusion strategy leverages the recognition advantages of each modality in specific situations to achieve multi-faceted information complementarity, significantly reducing the misclassification rate caused by missing or distorted sample features in a single classifier. Furthermore, combined with the dynamic decision boundary constructed in the superior claim, the ability to handle edge samples is further enhanced, improving the accuracy of the malware identification system in complex environments.

[0124] In this embodiment, the accuracy and generalization ability of identification in unknown malicious sample environments can be effectively improved. By extracting and fusing multimodal features such as static, dynamic, and API call sequences, the model has a more comprehensive perception ability of malicious software samples. The introduction of a triplet loss function realizes feature alignment and discriminative enhancement, improving inter-class separability. The radial basis function support vector machine (SVC) is used to construct the decision boundary, and the classification results are fused with a multi-stream threshold network, thereby achieving accurate identification of known categories and effective rejection of unknown categories, improving identification accuracy, and significantly enhancing detection performance and security defense capabilities in open set environments.

[0125] In this embodiment, the EMSTNet (Multi-Stream Thresholding Network) model was analyzed by comparing the performance of models with different embedding dimensions (128 / 256 / 512 units), and 256 units were selected as the optimal configuration. This dimension effectively balances feature discrimination ability while maintaining computational efficiency. The network architecture adopts a fully connected sequence structure with 512 to 256 layers, and its classification performance is verified to be superior to other configurations through layer-by-layer dimensionality reduction testing. The training process adopts a standardized protocol, with an initial learning rate of 0.0001 and 200 complete training cycles. For open set detection requirements, a radial basis function support vector classifier (RBF-SVCs) based on adaptive threshold of class center was constructed, with the threshold set to 2.0 times the standard deviation of the class center. All experiments were performed in a strictly controlled hardware environment, using an NVIDIA A30 graphics processor and an Intel Xeon Silver 4310 CPU, and maintaining software version consistency to ensure computational reproducibility.

[0126] Please refer to Figure 7 , Figure 7 This is a schematic diagram of the feature space for t-SNE two-dimensional visualization provided in an embodiment of the present invention, wherein, Figure 7 (a) represents the original feature distribution. Figure 7 (b) Optimize the distribution after metric learning.

[0127] In this embodiment, to intuitively verify the quality of the feature embedding space optimized by the metric learning module and triplet loss, the 256-dimensional fused feature vectors of all test samples are first input into the t-SNE algorithm for dimensionality reduction and visualization, and a dot plot is drawn on a two-dimensional plane. Samples of the same known family show high clustering on the t-SNE plot, with clear inter-cluster spacing, fully demonstrating the effective enhancement of intra-class aggregation and inter-class separability by metric learning. Subsequently, the classification results of the integrated multi-stream thresholding network (EMSTNet) on the same test set are organized into a confusion matrix; the values ​​of the main diagonal units of the matrix are highly concentrated, indicating excellent classification accuracy for each known family, while the misclassification rate at the off-diagonal positions is extremely low, verifying that the proposed scheme maintains high classification accuracy while also significantly reducing the rejection rate of unknown categories. This visualization and confusion matrix analysis jointly demonstrate the discriminative performance and robustness of the proposed method in open set scenarios.

[0128] In this embodiment, the EMSTNet model focuses on systematically optimizing two key hyperparameters: the triplet loss boundary parameter γ and the composite scoring tradeoff parameter λ. A discrete parameter space γ,λ∈[0.2,1.0] (step size 0.2) is constructed on the BODAS benchmark dataset, and a grid search is performed through 50 training cycles. Figure 7As shown, the convergence trajectories of test accuracy and loss value under different γ / λ combinations indicate that the combination of latent space separation parameter γ = 0.6 and decision boundary optimization parameter λ = 0.8 has equivalent importance in classifier optimization. Finally, empirical results show that this parameter configuration can optimally balance feature space separation degree and decision boundary discriminative power.

[0129] Please refer to Figure 8 , Figure 8 This is a schematic diagram of the prediction confusion matrix of each classifier provided in the embodiments of the present invention, wherein: Figure 8 (a) For static classifiers, Figure 8 (b) For dynamic classifiers, Figure 8 (c) For API function sequence classifier and Figure 8 (d) is the confusion matrix of the ensemble classifier.

[0130] The EMSTNet model, through Figure 6 The framework shown implements the detection process. After the feature vector is optimized by the metric learning module, it is input into the multi-stream classifier. The dynamic dual threshold adjustment module optimizes the confidence and boundary thresholds based on the validation set data. Finally, the result is generated through soft voting and dual decision rules. Figure 7 The comparison shows that metric learning significantly improves intra-class aggregation and inter-class separability in the feature space; Figure 8 The confusion matrix further demonstrates that the ensemble model exhibits a high concentration of values ​​on the main diagonal and a significant reduction in off-diagonal misclassification density, showcasing its advantages in reducing misclassification and improving unknown detection capabilities. Experimental results show that, in scenarios with 20% openness, the method achieves an unknown family detection rate (DetAcc) of 92.72% and a known family classification accuracy (ClsAcc) of 95.53%, making it suitable for scenarios such as enterprise endpoint protection, cloud security platforms, and IoT devices. This invention optimizes feature distribution through metric learning, balances classification and detection performance through dynamic thresholding, and enhances robustness through multi-stream ensemble, achieving high-precision malware identification in open-set scenarios. Its technical solution is adaptable to real-time detection requirements, providing effective support for proactive defense against unknown threats.

[0131] Please refer to Figure 9 , Figure 9 A schematic diagram of a malware open set identification device provided in an embodiment of the present invention includes: a feature extraction module 901, a feature fusion module 902, a decision adjustment module 903, and an identification module 904;

[0132] The feature extraction module 901 is used to acquire the sample data to be detected, perform feature extraction on the sample data to be detected, and acquire the multimodal features of each sample to be detected in the sample data to be detected.

[0133] The feature fusion module 902 is used to perform alignment and weighted fusion processing on the multimodal features of each sample to be detected based on a preset triplet loss function, so as to obtain the fused feature vector of each sample to be detected.

[0134] The decision adjustment module 903 is used to construct a radial basis function support vector classifier, and to identify the fused feature vector based on the radial basis function support vector classifier to determine the decision boundary;

[0135] The recognition module 904 is used to construct a multi-stream threshold network based on a preset static stream classifier, a preset dynamic classifier, and a preset API stream classifier, to recognize the multimodal features based on the multi-stream threshold network, to obtain the category label of each sample to be detected, and to generate the detection result of the sample data to be detected based on the decision boundary and the category label.

[0136] In this embodiment, the feature extraction module is used to obtain the sample data to be detected, which includes executable files and sandbox logs of several samples to be detected. The process of acquiring the sample data to be detected and performing feature extraction on the sample data to obtain the multimodal features of each sample to be detected in the sample data includes:

[0137] Obtain the executable file of the sample to be detected, extract features from the executable file, and obtain the static features of the sample to be detected;

[0138] Feature extraction is performed on the sandbox logs to obtain the running behavior data of the sample to be detected, and dynamic features of the sample to be detected are generated based on the running behavior data;

[0139] Feature extraction is performed on the sandbox logs to obtain the API call sequence features of the sample to be detected;

[0140] The static features, dynamic features, and API call sequence features constitute the multimodal features of the sample to be detected.

[0141] In this embodiment, the feature fusion module is used to perform alignment and weighted fusion processing on the multimodal features of each sample to be detected based on a preset triplet loss function, to obtain the fused feature vector of each sample to be detected, including:

[0142] The static features, dynamic features, and API call sequences of each sample to be detected are normalized to obtain normalized multimodal features.

[0143] The normalized multimodal features are mapped based on a preset metric learning network to generate embedding vectors for each sample to be detected.

[0144] The embedding vector is optimized based on a preset triplet loss function to generate a fused feature vector for each sample to be detected.

[0145] In this embodiment, the decision adjustment module is used to construct a radial basis function support vector classifier, and to identify the fused feature vector based on the radial basis function support vector classifier to determine the decision boundary, including:

[0146] Constructing a dynamic classifier based on radial basis function support vector classifier;

[0147] The fused feature vector is input into a pre-trained radial basis function support vector classifier to obtain the classification confidence and classification label of each sample to be detected;

[0148] Calculate the geometric distance from the sample to be detected to the classification label based on the classification label;

[0149] The decision boundary is determined based on the classification confidence level and geometric distance.

[0150] In this embodiment, the recognition module is used to construct a multi-stream threshold network based on a preset static stream classifier, a preset dynamic classifier, and a preset API stream classifier; to recognize the multimodal features based on the multi-stream threshold network; to obtain the category labels of each sample to be detected; and to generate the detection result of the sample data to be detected based on the decision boundary and the category labels, including:

[0151] Based on the preset static flow classifier, preset dynamic classifier and preset API flow classifier, static features, dynamic features and API call sequence features are identified respectively, and static identification results, dynamic identification results and API identification results are obtained;

[0152] The static recognition results, dynamic recognition results, and API recognition results are weighted and fused based on a preset soft voting mechanism to obtain the category label of each sample to be detected, and the detection result of the sample data to be detected is generated based on the decision boundary and the category label.

[0153] In this embodiment of the invention, a multi-device access platform processing device is also provided, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the above-described multi-device access platform processing method.

[0154] In this embodiment of the invention, a computer-readable storage medium is also provided, which includes a stored computer program, wherein the computer program controls the device where the computer-readable storage medium is located to execute the above-described multi-device access platform processing method when it is running.

[0155] For example, a computer program can be divided into one or more modules, one or more of which are stored in memory and executed by a processor to perform the present invention. The one or more modules can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in a multi-device access platform processing device.

[0156] The multi-device access platform processing device can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. The multi-device access platform processing device may include, but is not limited to, a processor, memory, and a display. Those skilled in the art will understand that the above components are merely examples of the multi-device access platform processing device and do not constitute a limitation on the multi-device access platform processing device. It may include more or fewer components than the specified components, or a combination of certain components, or different components. For example, the multi-device access platform processing device may also include input / output devices, network access devices, buses, etc.

[0157] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the multi-device access platform processing device, connecting all parts of the multi-device access platform processing device through various interfaces and lines.

[0158] Memory can be used to store computer programs and / or modules. The processor, by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory, implements various functions of the multi-device access platform processing device. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one application program required for a function (such as sound playback, text conversion, etc.); the data storage area can store data created based on the use of the mobile phone (such as audio data, text message data, etc.). Furthermore, memory can include high-speed random access memory, and can also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital cards (SD cards), flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.

[0159] In this invention, modules for processing multi-device access platforms, if implemented as software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. Computer-readable media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. Those skilled in the art can understand and implement this invention without any inventive effort.

[0160] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A method for identifying open sets of malicious software, characterized in that, include: Acquire the sample data to be detected, extract features from the sample data to be detected, and obtain the multimodal features of each sample to be detected in the sample data to be detected; Based on the preset triplet loss function, the multimodal features of each sample to be detected are aligned and weighted and fused to obtain the fused feature vector of each sample to be detected. A radial basis function support vector classifier is constructed, and the fused feature vector is identified based on the radial basis function support vector classifier to determine the decision boundary; A multi-stream threshold network is constructed based on a preset static stream classifier, a preset dynamic stream classifier, and a preset API stream classifier. The multi-modal features are identified based on the multi-stream threshold network to obtain the category label of each sample to be detected. The detection result of the sample data to be detected is generated based on the decision boundary and the category label.

2. The method for identifying open sets of malicious software as described in claim 1, characterized in that, The sample data to be tested includes executable files and sandbox logs of several samples to be tested; the process of acquiring the sample data to be tested, performing feature extraction on the sample data to be tested, and obtaining the multimodal features of each sample to be tested in the sample data to be tested includes: Obtain the executable file of the sample to be detected, extract features from the executable file, and obtain the static features of the sample to be detected; Feature extraction is performed on the sandbox logs to obtain the running behavior data of the sample to be detected, and dynamic features of the sample to be detected are generated based on the running behavior data; Feature extraction is performed on the sandbox logs to obtain the API call sequence features of the sample to be detected; The static features, dynamic features, and API call sequence features constitute the multimodal features of the sample to be detected.

3. The method for identifying open sets of malicious software as described in claim 2, characterized in that, The pre-defined triplet loss function is used to align and weight the multimodal features of each sample to be detected, thereby obtaining the fused feature vector of each sample to be detected, including: The static features, dynamic features, and API call sequences of each sample to be detected are normalized to obtain normalized multimodal features. The normalized multimodal features are mapped based on a preset metric learning network to generate embedding vectors for each sample to be detected. The embedding vector is optimized based on a preset triplet loss function to generate a fused feature vector for each sample to be detected.

4. The method for identifying open sets of malicious software as described in claim 3, characterized in that, The construction of a radial basis function support vector classifier, and the identification of the fused feature vector based on the radial basis function support vector classifier to determine the decision boundary, includes: Constructing a dynamic classifier based on radial basis function support vector classifier; The fused feature vector is input into a pre-trained radial basis function support vector classifier to obtain the classification confidence and classification label of each sample to be detected; Calculate the geometric distance from the sample to be detected to the classification label based on the classification label; The decision boundary is determined based on the classification confidence level and geometric distance.

5. The method for identifying open sets of malicious software as described in claim 4, characterized in that, The process involves constructing a multi-stream threshold network based on a preset static stream classifier, a preset dynamic stream classifier, and a preset API stream classifier. This multi-stream threshold network is used to identify the multimodal features, obtain the category labels for each sample to be detected, and generate detection results for the sample data based on the decision boundary and the category labels. This includes: Based on the preset static flow classifier, preset dynamic classifier and preset API flow classifier, static features, dynamic features and API call sequence features are identified respectively, and static identification results, dynamic identification results and API identification results are obtained; The static recognition results, dynamic recognition results, and API recognition results are weighted and fused based on a preset soft voting mechanism to obtain the category label of each sample to be detected, and the detection result of the sample data to be detected is generated based on the decision boundary and the category label.

6. A malware open set identification device, characterized in that, include: The module comprises a feature extraction module, a feature fusion module, a decision adjustment module, and a recognition module. The feature extraction module is used to acquire the sample data to be detected, perform feature extraction on the sample data to be detected, and acquire the multimodal features of each sample to be detected in the sample data to be detected. The feature fusion module is used to perform alignment and weighted fusion processing on the multimodal features of each sample to be detected based on a preset triplet loss function, so as to obtain the fused feature vector of each sample to be detected. The decision adjustment module is used to construct a radial basis function support vector classifier, and to identify the fused feature vector based on the radial basis function support vector classifier to determine the decision boundary; The recognition module is used to construct a multi-stream threshold network based on a preset static stream classifier, a preset dynamic classifier, and a preset API stream classifier, recognize the multimodal features based on the multi-stream threshold network, obtain the category label of each sample to be detected, and generate the detection result of the sample data to be detected based on the decision boundary and the category label.

7. The malware open set identification device as described in claim 6, characterized in that, The feature extraction module is used to obtain the sample data to be detected, which includes executable files and sandbox logs of several samples to be detected. The process of acquiring the sample data to be detected and performing feature extraction on the sample data to obtain the multimodal features of each sample to be detected in the sample data includes: Obtain the executable file of the sample to be detected, extract features from the executable file, and obtain the static features of the sample to be detected; Feature extraction is performed on the sandbox logs to obtain the running behavior data of the sample to be detected, and dynamic features of the sample to be detected are generated based on the running behavior data; Feature extraction is performed on the sandbox logs to obtain the API call sequence features of the sample to be detected; The static features, dynamic features, and API call sequence features constitute the multimodal features of the sample to be detected.

8. The malware open set identification device as described in claim 7, characterized in that, The feature fusion module is used to perform alignment and weighted fusion processing on the multimodal features of each sample to be detected based on a preset triplet loss function, to obtain the fused feature vector of each sample to be detected, including: The static features, dynamic features, and API call sequences of each sample to be detected are normalized to obtain normalized multimodal features. The normalized multimodal features are mapped based on a preset metric learning network to generate embedding vectors for each sample to be detected. The embedding vector is optimized based on a preset triplet loss function to generate a fused feature vector for each sample to be detected.

9. A malware open set identification device as described in claim 8, characterized in that, The decision adjustment module is used to construct a radial basis function support vector classifier, identify the fused feature vector based on the radial basis function support vector classifier, and determine the decision boundary, including: Constructing a dynamic classifier based on radial basis function support vector classifier; The fused feature vector is input into a pre-trained radial basis function support vector classifier to obtain the classification confidence and classification label of each sample to be detected; Calculate the geometric distance from the sample to be detected to the classification label based on the classification label; The decision boundary is determined based on the classification confidence level and geometric distance.

10. The malware open set identification device as described in claim 9, characterized in that, The recognition module is used to construct a multi-stream threshold network based on a preset static stream classifier, a preset dynamic classifier, and a preset API stream classifier; to recognize the multimodal features based on the multi-stream threshold network; to obtain the category label of each sample to be detected; and to generate the detection result of the sample data to be detected based on the decision boundary and the category label, including: Based on the preset static flow classifier, preset dynamic classifier and preset API flow classifier, static features, dynamic features and API call sequence features are identified respectively, and static identification results, dynamic identification results and API identification results are obtained; The static recognition results, dynamic recognition results, and API recognition results are weighted and fused based on a preset soft voting mechanism to obtain the category label of each sample to be detected, and the detection result of the sample data to be detected is generated based on the decision boundary and the category label.