File-free attack detection method, system and equipment based on multi-view behavior modeling and frequency domain enhanced contrast learning, and medium
By employing multi-view behavior modeling and frequency domain enhanced contrastive learning, this method addresses the shortcomings of fileless attack detection technology in terms of concealment, dynamism, and adversarial capabilities, achieving efficient, accurate identification and adaptive detection of memory-resident attack behaviors.
Patent Information
- Application Number
- CN202510821313.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-11-14
AI Technical Summary
Existing fileless attack detection technologies are significantly inadequate in dealing with the stealth, dynamism, polymorphism, and adversarial nature of attacks, and are unable to effectively identify memory-resident attack behaviors.
We employ a multi-view behavior modeling and frequency domain enhanced contrastive learning approach. By collecting multi-source behavior data during system operation, we construct time series and behavior graphs, extract features using self-attention mechanisms and graph structure neural networks, generate frequency domain representations through fast Fourier transform, and construct a joint contrastive learning loss function for consistency detection.
It achieves efficient and accurate identification of fileless attacks, possesses unsupervised learning, high robustness and strong generalization ability, and can adapt to the ever-evolving adversarial techniques of attackers.
Smart Images

Figure CN120956440A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network security technology, specifically to a method, system, device, and medium for detecting fileless attacks based on multi-view behavior modeling and frequency domain enhanced contrastive learning. Background Technology
[0002] With the continuous evolution of cyberattack techniques and the increasing sophistication of attackers' capabilities, traditional defense mechanisms relying on static signatures and file persistence are facing unprecedented challenges. Attackers are increasingly inclined to use more covert and evasive methods to launch attacks, aiming to bypass the detection of mainstream endpoint security products (such as traditional antivirus software and intrusion detection systems). Among these, fileless attacks, as a typical advanced persistent threat (APT) method, are gradually replacing traditional malware forms that rely on disk files, becoming one of the most challenging problems in the current cybersecurity field. The core characteristic of this type of attack is that it does not rely on writing malicious payloads to disk, but instead directly utilizes legitimate system processes, built-in tools (Living Off The Land Binaries and Scripts, LOLBAS), or legitimate APIs to dynamically load and execute malicious code in memory, thereby effectively circumventing traditional protection mechanisms based on file signatures, hash matching, or disk I / O monitoring.
[0003] Common fileless attack methods include, but are not limited to: using scripting languages to execute malicious commands: for example, using built-in scripting engines such as PowerShell, WMI (Windows Management Instrumentation), JavaScript, or VBScript to directly download, decrypt, and execute malicious code in memory, or call system APIs to perform sensitive operations, such as stealing credentials (Invoke-Mimikatz), establishing C2 communication, or persistence.
[0004] Memory injection and process hijacking: Through complex techniques such as DLL injection, reflective PE loading, process hole (process hole / RunPE), and Asynchronous Procedure Call Injection, malicious code is injected into the memory space of legitimate processes (such as explorer.exe, svchost.exe, etc.), making it behave like a legitimate process, thereby executing malicious operations and exploiting the privileges of the host process.
[0005] API abuse and remote thread creation: Using APIs such as CreateRemoteThread, NtCreateThreadEx, and QueueUserAPC to create and execute threads in remote processes, or to pass malicious payloads between processes through memory-mapped files, shared memory, etc.
[0006] Exploiting kernel vulnerabilities to achieve non-persistent memory-resident operations: By exploiting vulnerabilities in the operating system kernel or drivers, malicious code can be directly manipulated under Ring0 privileges, and usually no trace is left on the hard drive, making it extremely stealthy.
[0007] Other covert techniques include: using .NET Assembly loading, COM hijacking, code injection into code holes (CodeCaves), or using various reflection techniques to execute assemblies directly in memory without writing to disk.
[0008] Because these attack chains mostly run in memory, do not generate disk files, and are highly dynamic, polymorphic, and have memory volatility (for example, malicious code can self-clean up after running in memory), as well as strong anti-forensic capabilities (memory data is lost after a reboot), their behavior is difficult to detect by traditional file scanning, static signature antivirus engines, or sandbox-based analysis mechanisms.
[0009] Existing fileless attack detection technologies can be mainly divided into the following three categories, but all of them have significant limitations:
[0010] Static detection methods based on rules or feature libraries, behavior detection methods based on sandbox environments, and sequence modeling methods based on traditional machine learning are all existing methods for detecting fileless attacks. However, these methods have significant shortcomings in addressing the stealth, dynamism, polymorphism, and adversarial nature of attacks. To effectively combat the severe security challenges posed by fileless attacks, there is an urgent need for a novel detection mechanism that can integrate multimodal behavioral data, deeply analyze the intrinsic structural dependencies of behavior, and innovatively introduce frequency domain feature enhancement. This mechanism should be able to capture deep patterns of behavioral events from multiple perspectives, thereby achieving efficient, accurate, and robust identification of fileless attack behaviors lurking in memory.
[0011] This invention focuses particularly on the detection of "fileless attacks." This technology achieves real-time, efficient, and accurate identification of memory-resident attacks in highly dynamic and covert environments by deeply modeling and analyzing multi-source behavioral data during system runtime, such as CPU instruction execution sequences, memory access patterns, inter-process communication, system calls, and API function calls. Its core lies in the innovative introduction of a multi-view behavioral modeling strategy, jointly considering the time-series patterns of behavioral events, their inherent contextual structure (such as call chains and data flows), and the characteristic changes of behavior in the frequency domain. Combined with a frequency-domain enhanced contrastive learning mechanism, this method can adaptively learn stable behavioral patterns from a large amount of normal behavioral data and identify abnormal behaviors that deviate from these patterns without prior knowledge of attack samples. The proposed method has advantages such as unsupervised learning, high robustness, and strong generalization ability, effectively addressing the ever-evolving adversarial techniques of attackers. Summary of the Invention
[0012] In view of the above-mentioned problems, the present invention is proposed.
[0013] Therefore, the technical problem solved by this invention is: how to address the significant shortcomings of existing fileless attack detection technologies in dealing with the concealment, dynamism, polymorphism, and adversarial nature of attacks, as well as the significant limitations of existing fileless attack detection technologies in understanding dynamic behavior, model robustness, and lack of frequency domain analysis.
[0014] To address the aforementioned technical problems, this invention provides the following technical solution: a fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning, comprising: collecting and encoding multi-source behavior data of the target system during operation, and performing unified encoding; dividing the multi-source behavior data into time series and constructing corresponding behavior graphs; inputting the divided time series into a self-attention mechanism neural network to extract the time domain representation of the behavior, and inputting the constructed behavior graph into a graph structure neural network to extract the structure representation; performing fast Fourier transform on the time domain representation and the structure representation respectively to generate a frequency domain representation; constructing a joint contrastive learning loss function, wherein the loss function includes a time domain loss that measures the difference between time domain representations using symmetric KL divergence, and a frequency domain loss that measures the difference between frequency domain representations using mean absolute error, and training a consistency detection model; in the inference stage, processing the sample to be detected through the consistency detection model, calculating the consistency deviation between multimodal representations, and determining whether it is a fileless attack behavior based on the consistency deviation and combined with an anomaly detection judgment mechanism.
[0015] As a preferred embodiment of the fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning described in this invention, the self-attention mechanism neural network includes modeling the time domain representation and using a multi-head self-attention structure to extract time-dependent and sequence semantic features.
[0016] As a preferred embodiment of the fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning described in this invention, the graph structure neural network includes modeling the structural representation, extracting structural dependencies and contextual features between nodes through the aggregation of node adjacency information, and generating the structural representation.
[0017] As a preferred embodiment of the documentless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning described in this invention, the joint contrastive learning loss function includes a time domain loss measured by symmetric KL divergence and a frequency domain loss measured by mean absolute error. The process of training the consistency detection model is optimized using the joint contrastive learning loss function.
[0018] As a preferred embodiment of the fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning described in this invention, the time domain representation refers to the representation generated after inputting the time series into a self-attention mechanism neural network. The self-attention mechanism neural network is a neural network based on a multi-head self-attention mechanism, which dynamically calculates the correlation between any two positions in the time series, extracts the temporal order, duration, and sequence expression features of behavioral events, and generates a representation vector that characterizes the temporal semantics of the time series.
[0019] This preferred solution uses a neural network based on a multi-head self-attention mechanism to model time series. It can dynamically identify the correlation between any two behavioral events without relying on fixed time window features, thereby accurately extracting key features that represent the time sequence and execution rhythm of behavioral events. Even when the attack behavior is hidden in the normal process and the order of behavior is disrupted or reconstructed, it can still maintain a stable temporal semantic extraction capability.
[0020] As a preferred embodiment of the fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning described in this invention, the structural representation refers to the representation generated after inputting the behavior graph into a graph structure neural network. The graph structure neural network performs feature aggregation and neighborhood modeling on the nodes in the behavior graph, learns the nonlinear structural dependencies between behavior events, and uses graph-level pooling to aggregate and generate a structural representation vector representing the overall structure of the behavior graph from all node representations.
[0021] This preferred solution inputs the behavioral graph into a graph structure neural network, and obtains the structural representation through feature aggregation between nodes and graph-level pooling. It can model the structural relationships between various types of edges such as function call chains, data flows, and control flows, thereby effectively expressing the dependencies under nonlinear behavioral paths. This helps to identify the structural linkages between behaviors that have not been written to disk in hidden paths such as process holes and DLL injection during the detection process.
[0022] As a preferred embodiment of the fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning described in this invention, the frequency domain representation refers to the representation generated by performing fast Fourier transform on the time domain representation and the structural representation respectively. The fast Fourier transform converts the time domain representation and the structural representation from time domain signals to frequency domain signals, obtaining the energy distribution at different frequencies, reflecting periodic fluctuations or frequency component changes. The frequency domain representation serves as the input to the frequency domain loss in the joint contrastive learning loss function.
[0023] This preferred scheme generates a frequency domain representation by performing fast Fourier transforms on the time domain representation and the structural representation respectively. This allows the model to further observe the changing trends of behavioral data from the frequency dimension, thereby identifying attack characteristics such as periodic anomalies (e.g., fixed-interval communication, periodic memory access) or frequency mutations (e.g., batch calls, intensive access), and enhancing the detection system's ability to perceive dynamic evasion behaviors.
[0024] This invention provides a fileless attack detection system based on multi-view behavior modeling and frequency domain enhanced contrastive learning.
[0025] To address the aforementioned technical problems, this invention provides the following technical solution: a fileless attack detection system based on multi-view behavior modeling and frequency domain enhanced contrastive learning, comprising: a data acquisition and encoding module, a behavior graph construction module, a feature extraction module, a frequency domain transformation module, a consistency modeling module, and an anomaly detection module; the data acquisition and encoding module is used to acquire and encode multi-source behavior data of the target system during operation, and perform unified encoding; the behavior graph construction module is used to divide the multi-source behavior data into time series and construct corresponding behavior graphs; the feature extraction module is used to input the divided time series into a self-attention mechanism neural network, extract the time domain representation of the behavior, and construct the... The behavioral graph is input into a graph structure neural network to extract structural representations. The frequency domain transformation module performs Fast Fourier Transform on the time domain representation and the structural representation respectively to generate a frequency domain representation. The consistency modeling module constructs a joint contrastive learning loss function, which includes a time domain loss that measures the difference between time domain representations using symmetric KL divergence and a frequency domain loss that measures the difference between frequency domain representations using mean absolute error, to train a consistency detection model. The anomaly detection module processes the samples to be detected through the consistency detection model during the inference phase, calculates the consistency deviation between multimodal representations, and determines whether it is a fileless attack based on the consistency deviation and an anomaly detection judgment mechanism.
[0026] The present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, characterized in that the processor executes the computer program to implement the steps of the fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning.
[0027] The present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the steps of the fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning.
[0028] The beneficial effects of this invention are: This invention does not rely on preset manual rules, signatures, or complex feature engineering. Through data-driven comparative learning, the model can adaptively learn universal normal patterns from massive amounts of normal behavior data, exhibiting stronger generalization ability and robustness against unknown attack variants and constantly evolving obfuscation and evasion techniques.
[0029] This invention innovatively integrates temporal patterns of behavior, contextual structural dependencies, and frequency features. It captures long-term temporal dependencies through a self-attention mechanism, reveals complex call chains and data flows through graph neural networks, and discovers hidden periodicity or energy perturbations through frequency domain analysis. This multi-view fusion modeling can more comprehensively and deeply characterize the multi-stage, multi-modal behavioral chains of fileless attacks, significantly improving detection accuracy.
[0030] This invention employs a contrastive learning paradigm, requiring only a large number of normal behavior samples for training, without the need for scarce and difficult-to-obtain malicious sample labels. This significantly lowers the barrier to model training, enabling it to better adapt to the challenges of collecting diverse and constantly updated attack samples in real-world environments.
[0031] This invention provides a general behavioral analysis framework that is not targeted at specific attack signatures or single attack methods. It can uniformly and effectively detect various types of fileless attack behaviors, such as those using PowerShell scripts, DLL injection, process holes, API abuse, and other memory-resident operations, thereby improving the universality and efficiency of the detection system. Attached Figure Description
[0032] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0033] Figure 1 The above is a flowchart of a fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning, provided as an embodiment of the present invention.
[0034] Figure 2 This is a flowchart illustrating a fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning, provided as an embodiment of the present invention.
[0035] Figure 3 The diagram shows a fileless attack detection model based on multi-view behavior modeling and frequency domain enhanced contrastive learning, which is provided as an embodiment of the present invention.
[0036] Figure 4 This is a block diagram of a fileless attack detection system based on multi-view behavior modeling and frequency domain enhanced contrastive learning, provided as an embodiment of the present invention. Detailed Implementation
[0037] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0038] Example 1, referring to Figure 1 This is one embodiment of the present invention, which provides a fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning, including:
[0039] S1. Collect and encode multi-source behavioral data of the target system during operation, and perform unified encoding.
[0040] S2. Divide the multi-source behavioral data into time series segments and construct the corresponding behavioral graphs.
[0041] S3. Input the segmented time series into a self-attention mechanism neural network to extract the temporal representation of the behavior, and input the constructed behavior graph into a graph structure neural network to extract the structural representation.
[0042] S4. Perform Fast Fourier Transform on the time domain representation and the structural representation respectively to generate the frequency domain representation;
[0043] S5. Construct a joint contrastive learning loss function, which includes a time domain loss that measures the difference between time domain representations using symmetric KL divergence and a frequency domain loss that measures the difference between frequency domain representations using mean absolute error, and train the consistency detection model.
[0044] S6. In the inference phase, the sample to be tested is processed by the consistency detection model, the consistency deviation between multimodal representations is calculated, and the consistency deviation is used in conjunction with the anomaly detection judgment mechanism to determine whether it is a fileless attack.
[0045] Existing fileless attack detection technologies can be mainly divided into the following three categories, but all of them have significant limitations:
[0046] Static detection methods based on rules or signature libraries rely on predefined attack patterns, malicious code hash features, specific API call sequences, or heuristic rules based on string matching. For example, monitoring common obfuscated keywords in PowerShell commands or import table features of specific DLLs. Its limitations include the vulnerability of these rules to rapidly evolving attacker variants, obfuscation (such as PowerShell obfuscation and code obfuscation), encryption strategies (such as AES-encrypted shellcode), and polymorphic and metamorphic techniques. These rules lack generalization ability and struggle to identify unknown or variant fileless attacks (i.e., high false negative rate). Attackers can easily bypass detection based on known features by modifying code, parameters, etc. Furthermore, these methods can also lead to a high false positive rate because some legitimate actions may occasionally violate general rules.
[0047] Behavioral detection methods based on sandbox environments: Suspicious programs are dynamically run in isolated virtual machines or simulation environments, and their runtime behavior (such as file operations, registry modifications, network connections, and system call sequences) is observed to determine the presence of malicious activity. Its limitations include: limited execution window: Sandboxes typically have preset execution time limits, which attackers can circumvent by delaying execution or waiting for specific environmental triggers (such as user interaction or network connection establishment); adversarial evasion techniques: Malware possesses the ability to recognize sandbox environments (such as detecting virtual machine instructions, specific memory layouts, virtual hardware information, and non-real user behavior), and once it detects itself in a sandbox, it stops its malicious behavior or exits directly, leading to detection failure; mismatched environment triggers: Sandbox environments cannot fully simulate real, complex production environments; certain attack behaviors require specific system states, network topologies, or user interactions to trigger, and the sandbox may not accurately reconstruct the entire attack chain; high resource consumption: Sandbox analysis requires simulating a complete operating system environment, resulting in high resource consumption, making it difficult to achieve large-scale, high-concurrency real-time host behavior monitoring, and it is unsuitable for edge devices.
[0048] Sequence modeling methods based on traditional machine learning: These methods typically use a single type of runtime data (such as system call sequences, CPU instruction sequences, API call frequencies, or instruction n-grams) as feature input, and utilize statistical methods (such as Hidden Markov Models, HMMs), traditional machine learning algorithms (such as Support Vector Machines, SVMs, and decision trees), or deep learning models (such as Recurrent Neural Networks, Long Short-Term Memory Networks, and LSTMs) for classification modeling or anomaly detection. Their limitations include: feature engineering constraints: they often rely on manually designed static features, lacking a deep understanding of the complex semantics of attack behaviors; and neglecting the intrinsic correlations of behaviors: these methods typically only focus on co-occurrence patterns or frequency statistics of behaviors in time series, often ignoring the long-term dependencies of behavioral events in time series, the deep semantics in the context structure, and the intrinsic correlations between multi-source data. For example, structured information such as inter-process communication, function call chains, and data flow is simply flattened, failing to capture the complex, multi-stage, cross-modal behavioral chain characteristics and dynamic evolution of fileless attacks. There is a lack of deep modeling capabilities for the evolution of attack behavior: relying solely on local sequence patterns or frequency statistics makes it difficult to distinguish between occasional abnormal behaviors in normal system activity and purposeful, organized fileless attack chains, thus making it susceptible to false positives and false negatives. There is also a single-modality limitation: most current research tends to focus on single-modality or single-dimensional data modeling (e.g., considering only instruction flows or system call sequences), neglecting the multi-perspective information contained in the temporal, structural, and frequency domains of behavioral events. This one-sided modeling approach cannot fully reveal the overall characteristics and subtle dynamic changes of fileless attack behavioral chains.
[0049] This invention belongs to the field of information security and network attack and defense technology, and specifically focuses on the detection and defense against advanced malicious behaviors in response to the increasingly severe threat landscape. It specifically involves key technical directions such as host security protection, runtime memory security detection, advanced persistent threat (APT) defense, threat hunting, and endpoint security response. As attackers' skills continue to improve, traditional signature-based detection methods are no longer sufficient; therefore, behavioral analysis and anomaly detection have become a research hotspot and mainstream direction in the current network security field.
[0050] This invention focuses particularly on the detection of the most challenging fileless attacks. The technology employs deep modeling and analysis of multi-source behavioral data during system runtime, such as CPU instruction execution sequences, memory access patterns, inter-process communication, system calls, and API function calls, to achieve real-time, efficient, and accurate identification of memory-resident attacks in highly dynamic and covert environments. Its core lies in the innovative introduction of a multi-view behavioral modeling strategy, jointly considering the time-series patterns of behavioral events, their inherent contextual structure (such as call chains and data flows), and the frequency domain's characteristic variations. Combined with a frequency domain-enhanced contrastive learning mechanism, this method can adaptively learn stable behavioral patterns from a large amount of normal behavioral data and identify anomalous behaviors deviating from these patterns without prior knowledge of attack samples. The proposed method possesses advantages such as unsupervised learning, high robustness, and strong generalization ability, effectively addressing the evolving adversarial techniques of attackers.
[0051] This invention can be widely applied to various computing environments, including but not limited to enterprise-level endpoint security protection (such as workstations, servers, and IoT devices), virtualization environment security monitoring (such as cloud hosts, containers, and hypervisors), cloud environment threat identification and response (such as serverless functions and threat detection in microservice architectures), and key scenarios requiring in-depth host behavior analysis and threat intelligence generation, providing core technical support for building a proactive defense system.
[0052] In modern computing environments, the detection of fileless attacks poses a significant challenge to traditional detection methods that rely on static signatures or single behavioral patterns due to their independence from file storage, highly dynamic processes, and memory residency. This invention constructs an unsupervised consistency detection process for multimodal behaviors by sequentially executing steps including multi-source behavioral data acquisition and encoding, time series partitioning and behavioral graph construction, parallel time and structural modeling, frequency domain transformation, contrastive learning modeling, and inference judgment. This process effectively captures the dynamic characteristics of behavior across time, structure, and frequency, enabling the identification and judgment of fileless attack behaviors, and possesses deployability and implementability.
[0053] Example 2, refer to Figure 2 and Figure 3 As an embodiment of the present invention, based on the previous embodiment, a fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning is provided, including:
[0054] In the embodiments of this application, step S1 involves collecting and encoding multi-source behavioral data of the target system during operation, and then performing unified encoding.
[0055] Collect multi-source behavioral data generated by the target system during runtime, including but not limited to memory access behavior, CPU instruction sequences, system calls and function call events, and encode them uniformly to form a multi-dimensional behavioral sequence containing information such as access address, operation type, call name, timestamp, and parameters.
[0056] To comprehensively capture the subtle traces left by fileless attacks during system runtime, this invention designs a dynamic, fine-grained behavior acquisition mechanism. This mechanism can monitor and record multi-source key behavioral information in the target system in real time, including but not limited to: CPU instruction sequences: such as opcodes, instruction execution flow, register state changes, and memory access patterns (read / write / execute permissions). Memory access behaviors: including memory region allocation, release, read / write operations, page protection attribute modifications (e.g., from read-write to execute), and stack activity. System call events: recording the name, parameters, return value, calling process / thread ID, and timestamp of system calls, revealing the interaction pattern between the program and the operating system. API function call events: capturing the parameters, return values, and call stack information of key API calls in user mode and kernel mode. By deploying a lightweight, low-overhead runtime data acquisition agent, the above raw events are captured in real time and uniformly encoded into structured multi-dimensional behavior vector sequences, serving as the basic input for subsequent modeling.
[0057] Specifically, a runtime data acquisition agent is deployed in the target system to monitor the following behavioral events in real time: CPU instruction sequences (such as opcode, register status); memory access behaviors (address, size, access type); API or system call events (call name, parameters, timestamp, thread ID, etc.). The above raw events are uniformly encoded into multi-dimensional behavior vectors, where each time step contains fields such as event type, behavior subject, behavior object, and context information, which constitute the raw input behavior sequence.
[0058] The encoding methods include: extracting access address digests and access types (read / write / execute) from memory events; extracting operation code streams and timing from instruction sequences; and embedding information such as call name, parameter types, return values, and call depth from function call records into a unified high-dimensional vector.
[0059] In one alternative implementation, the acquisition agent intercepts all critical syscall events, including process creation, thread scheduling, file handle operations, and dynamic library loading, through a system call hook mechanism. It records the call name, call parameters, return value, and call time to characterize the behavior execution path. It also captures the call stack depth, call frequency, and parameter type of user-mode API calls through instrumentation, and combines these with system calls to form behavior sequence fragments.
[0060] In another alternative implementation, behavior acquisition further integrates memory allocation and release event channels, identifies anonymous mapping, heap segment growth and memory protection attribute change operations, records the timing of their occurrence, memory range, mapping type and operation source pointer; and combines thread-level events such as priority switching, interrupt recovery and other micro-scheduling behaviors as supplementary information to enhance the ability to respond to micro-behavioral mutations.
[0061] This invention integrates key behavioral information from multiple dimensions, including CPU instruction streams, memory access, system calls, and API calls, and combines this with fine-grained multi-channel acquisition methods such as page attribute changes, call stack information, and thread scheduling characteristics. This enables a more comprehensive recording of the target system's dynamic state during runtime, effectively reconstructing the latent behavioral chain during the attack process. Simultaneously, by using a unified encoding mechanism, heterogeneous behaviors are standardized into structured vector representations, providing a stable and consistent data foundation for subsequent multimodal joint modeling.
[0062] In this embodiment of the application, step S2 involves dividing the multi-source behavioral data into time series segments and constructing corresponding behavioral graphs.
[0063] Specifically, multi-source behavioral data is divided into multiple patches using a time-sliding window, and a corresponding function call graph or behavioral dependency graph is constructed for each patch. The nodes of the graph represent behavioral events or function calls, and the edges represent temporal order, data dependencies, or control flow relationships. The fileless attack detection process is as follows: Figure 2 As shown.
[0064] Patch partitioning and behavior graph construction. To more effectively process long sequence data and capture the correlation between local and global data, the original behavior sequence is partitioned, and a structured graph is constructed for each partition.
[0065] Time series patching: The original input behavior sequence S is divided into multiple overlapping or non-overlapping subsequences P1, P2, ..., P by a sliding window with a fixed window length (e.g., 500 or 1000 behavior events) or an adaptive window length (e.g., based on process lifecycle, function execution integrity, or specific event markers). M Each P i A patch represents a snapshot of behavior over a specific time period. This segmentation helps the model focus on local behavioral patterns while maintaining contextual continuity through a sliding window mechanism.
[0066] Behavior graph construction: For each partitioned P iConstruct a corresponding behavioral dependency graph G = (V, E). Graph Nodes (V): Nodes in the graph represent atomic behavioral events or key entities occurring in the patch, such as: a specific API function call instance, a memory read / write / execution operation, a CPU instruction block (e.g., a basic block), or important system events (e.g., process creation, module loading). Nodes can carry their corresponding encoded vectors as initial features. Graph Edges (E): Edges in the graph represent relationships or dependencies between nodes. Edges can be constructed based on the following relationships: Temporal sequence: Event A occurs before event B. Control flow: Function call relationships (caller-callee), instruction jumps, parent-child relationships between processes / threads. Data dependency: The output of one event (e.g., return value, memory written) serves as the input of another event. Memory region association: Multiple events access or operate on the same memory region. Heterogeneous Graph Modeling: To more comprehensively capture complex behaviors, this invention supports the construction of heterogeneous graphs. This means that the graph can contain different types of nodes (e.g., API nodes, memory access nodes, instruction nodes) and different types of edges (e.g., call, access, data flow, control flow). Different types of nodes and edges can have their own feature vectors or relation weights, enabling graph-structured neural networks to model multimodal interactions in greater detail.
[0067] Behavioral graph construction methods include: using function calls or system behavior events as graph nodes; using temporal sequence, control dependency, data flow, etc. as the basis for edge construction; and supporting heterogeneous graph modeling, meaning that the types of nodes and edge types may not be completely consistent.
[0068] In one alternative implementation, the behavior graph is modeled using a heterogeneous graph structure. Node types include API nodes, memory access nodes, instruction nodes, etc., and edge types include control flow edges, data flow edges, and access edges. For example, for the VirtualAlloc→WriteProcessMemory→CreateRemoteThread call chain within the same patch, function call order edges and shared memory region access edges are constructed respectively to express multiple behavioral relationships across nodes.
[0069] In another alternative implementation, memory space location and object-level data access information are introduced during the behavior graph construction process. For example, multiple accesses to the same memory page will be aggregated to construct access set edges. At the same time, parent-child thread edges are introduced for child thread behavior nodes within the same process to improve the behavior graph's ability to express behavioral structures under multi-threaded context switching.
[0070] This invention constructs a unified and scalable behavioral graph structure, combined with a node and edge type classification method, which can clearly express the control, invocation, and data relationships between multi-source behavioral events within a single time segment, providing a well-structured and semantically clear input foundation for subsequent graph structure modeling.
[0071] In this embodiment of the application, in step S3, the segmented time series is input into a self-attention mechanism neural network to extract the temporal representation of the behavior, and the constructed behavior graph is input into a graph structure neural network to extract the structural representation.
[0072] Self-attention neural networks model time-domain representations and employ multi-head self-attention structures to extract time-dependent and sequence semantic features.
[0073] Graph neural networks model structural representations by extracting structural dependencies and contextual features between nodes through the aggregation of node adjacency information, thereby generating structural representations.
[0074] The time-domain representation refers to the representation generated after inputting a time series into a self-attention mechanism neural network. The self-attention mechanism neural network is a neural network based on a multi-head self-attention mechanism. It dynamically calculates the correlation between any two positions in the time series, extracts the temporal order, duration, and sequence expression features of behavioral events, and generates a representation vector that represents the temporal semantics of the time series.
[0075] Structural representation refers to the representation generated after inputting the behavior graph into a graph structure neural network. The graph structure neural network performs feature aggregation and neighborhood modeling on the nodes in the behavior graph, learns the non-linear structural dependencies between behavior events, and uses graph-level pooling to aggregate from all node representations to generate a structural representation vector that represents the overall structure of the behavior graph.
[0076] Specifically, the segmented time series patches are input into a neural network based on a self-attention mechanism to extract time dependencies and sequence semantic features, generating a time-domain representation Z. T The constructed behavioral graph is input into a graph structure neural network to extract structural dependencies and contextual features between nodes, generating a structural representation Z. G ;
[0077] To overcome the limitations of single-modal modeling, this invention innovatively introduces a multimodal joint modeling strategy to characterize complex behaviors from different perspectives.
[0078] Temporal pattern modeling (based on a self-attention neural network): The original input sequence of behaviors is first patched according to a fixed or adaptive window length, decomposing the long sequence into multiple subsequences with local temporal context. These patched sequences are then fed into a self-attention neural network to extract contextual dependencies and long-term behavioral patterns across time steps. The self-attention mechanism dynamically calculates the correlation between any two positions in the sequence, effectively capturing the temporal order, duration, and inherent sequence representation of behavioral events, generating a representation vector Z with rich temporal semantics. T .
[0079] Context structure modeling (based on graph-structured neural networks): For each behavior patch, we further construct a behavior dependency graph G = (V, E) based on its internal call relationships, data flow, process / thread interaction, memory access logic, etc. Here, node V represents discrete behavior events (such as API calls, memory operations, process creation), and edge E represents the logical relationships between these events, such as temporal sequence, control dependencies, parameter passing, parent-child relationships, and memory sharing. This invention supports the construction of heterogeneous graphs, where nodes can be categorized into API nodes, memory nodes, instruction nodes, etc., and edges also have different types. The constructed behavior graph is then input into a graph-structured neural network to perform feature aggregation and neighborhood modeling on the nodes, in order to learn a graph-structured representation vector Z that reflects the behavior chain and context structure. G This graph modeling approach can deeply explore the complex semantic relationships between behavioral events and capture structured information that is difficult to achieve with traditional sequence models.
[0080] Temporal dependencies and sequence semantic features are extracted from behavioral sequence patches. Each partitioned temporal sequence patch is input into a neural network based on a multi-head self-attention mechanism. This mechanism allows the model to compute the correlation between any two positions in the sequence in parallel at different points of interest. It effectively captures contextual dependencies across time steps, whether short-term local patterns or long-term, non-linear behavioral chains. The model can learn: the temporal evolution of behavioral patterns: recognizing the typical order and rhythm of behavioral events; long-distance correlations between behavioral events: overcoming the gradient vanishing / exploding problem in traditional recurrent neural networks (RNNs) and capturing key causal relationships in attack chains that may span a large number of unrelated events; and the overall semantics of the sequence: generating an abstract, high-dimensional temporal domain representation vector Z for the entire patch. T Z represents the overall temporal semantics of the behavioral sequence within that time window. T =Transformer({x1,x2,...,x...) n}), where x i It is the behavior vector of the i-th patch.
[0081] Graph Structured Neural Networks (GNNs) are used to extract structural dependencies and contextual features between nodes from a constructed behavior graph. The behavior graph constructed for each patch is input into the GNN, which iteratively updates the node representation by aggregating the neighbor information of each node. This means that each node learns the features and structural context of its direct and indirect neighbors. For heterogeneous graphs, the GNN can employ a message-passing mechanism, performing different aggregation operations based on edge and node types to extract structural dependencies and contextual features. The GNN effectively captures non-linear structural dependencies between behavioral events, such as the depth and breadth of function call chains, the complete path of data flow, and the temporal dependencies formed by multiple unrelated APIs in memory. Through multi-layer GNN propagation, the model obtains an abstract, high-dimensional structural representation vector Z for each behavior graph. G Z G It can be obtained by aggregating from all node representations through graph-level pooling (such as global average pooling, Set2Set) or the Readout function, representing the overall structural semantics of the graph. Z G =GNN(G=(V,E)). GNN's ability to model graph structures makes it more robust against complex evasion techniques such as obfuscation, code distortion, DLL injection, and process holes, because it focuses on the intrinsic relationships between behaviors rather than simple sequential features.
[0082] Graph-based neural networks offer better feature capture, handle complex dependency structures, and exhibit robustness against evasion attacks.
[0083] In the embodiments of this application, in step S4, fast Fourier transforms are performed on the time domain representation and the structural representation respectively to generate the frequency domain representation.
[0084] Frequency domain representation refers to the representation generated by performing Fast Fourier Transform (FFT) on the time domain representation and the structural representation respectively. The FFT converts the time domain representation and the structural representation from time domain signals to frequency domain signals, obtaining the energy distribution at different frequencies, reflecting periodic fluctuations or frequency component changes. The frequency domain representation serves as the input to the frequency domain loss in the joint contrastive learning loss function.
[0085] Specifically, for Z T With Z G Perform a fast Fourier transform to obtain its corresponding frequency domain representation F. T and F G .
[0086] To further enhance the model's ability to perceive behavioral anomalies, this invention introduces frequency domain analysis. Traditional time domain analysis is susceptible to noise and obfuscation, while frequency domain analysis can reveal the periodicity, suddenness, or anomalous energy distribution hidden in behavior. This invention applies joint multimodal representation (or its sub-representation Z) to... T With Z G Applying the Fast Fourier Transform to map it from the time domain to the frequency domain, we obtain the corresponding frequency domain representation F. T and F G These frequency domain features can capture the energy distribution, periodic fluctuations, and sudden frequency component disturbances of behavior at different frequencies, such as periodic beacon communication of malicious payloads, abnormal CPU instruction execution frequencies, or anomalous enhancement of high-frequency components in memory read / write patterns. These frequency domain features provide a complementary and robust perspective for anomaly detection.
[0087] The Fast Fourier Transform (FFT) is introduced to map the abstract representations of the time and structural domains to the frequency domain, capturing periodic or anomalous frequency perturbations in behavior. The generated time-domain representation Z is then analyzed. T and the generated structural domain representation Z G Applying the Fast Fourier Transform (FFT). The FFT transforms a signal from the time domain to the frequency domain, revealing its energy distribution at different frequencies. In behavioral analysis, it can help us discover: periodic behavior: for example, certain malicious operations may repeat at a fixed frequency (such as data theft). Anomalous frequency perturbations: normal behavioral patterns may have stable frequency characteristics, while attack behavior may introduce new, unusual frequency components or disrupt the original frequency distribution. Although Z... T and Z G Since it is a high-dimensional vector, we can treat it as a high-dimensional signal, perform a Fourier transform on it, and obtain its energy distribution in different frequency dimensions.
[0088] Frequency domain representation generation: After FFT processing, F T =FFT(Z) T ),F G =FFT(Z) G ), to obtain its corresponding frequency domain representation F T and F G These two frequency domain representations can provide information that complements the time and structural domains, offering a richer perspective for subsequent comparative learning.
[0089] In this embodiment of the application, step S5 involves constructing a joint contrastive learning loss function, which includes a time domain loss that measures the difference between time domain representations using symmetric KL divergence and a frequency domain loss that measures the difference between frequency domain representations using mean absolute error, and training a consistency detection model.
[0090] The joint contrastive learning loss function includes a time-domain loss measured by symmetric KL divergence and a frequency-domain loss measured by mean absolute error. The process of training the consistency detection model is optimized using the joint contrastive learning loss function.
[0091] This invention employs contrastive learning as its core training paradigm, aiming to address the scarcity of undocumented attack samples and the difficulty in obtaining labels, thereby achieving unsupervised or self-supervised learning. Its core idea is to utilize only massive amounts of normal behavior data during the training phase, maximizing the consistency of the same normal behavior across different modalities / views (e.g., frequency domain transformations of temporal view representations and graph structure view representations), while minimizing the consistency between different behaviors (or different randomly augmented views of the same behavior). Specifically, this invention constructs a temporal and frequency domain consistency contrastive loss. This loss function encourages the model to learn a compact representation space, allowing for multiple view representations of normal behavior (such as F...) to... T and F G Within this space, they are close to each other, forming a consistent distribution of normal behavioral patterns. By minimizing the Kullback-Leibler divergence or other contrastive loss functions, the model is trained to identify and memorize the intrinsic correlations and consistency boundaries of normal behavior in the time, structure, and frequency domains, without any malicious samples or artificial labels.
[0092] Specifically, a joint contrastive learning loss function is constructed, which aims to minimize the inconsistency between the time-domain representation and the structural representation in the time and frequency domains. The loss function includes: a time-domain loss that measures the difference between the time-domain representation and the structural representation, and a frequency-domain loss that measures the difference between the frequency-domain representation of the time-domain representation and the frequency-domain representation of the structural representation; wherein, the time-domain loss is measured by symmetric KL divergence, and the frequency-domain loss is measured by MAE.
[0093] A consistency detection model is trained under unsupervised conditions using a loss function, learning the distribution boundary of behavioral consistency based solely on normal samples.
[0094] A joint contrastive learning loss function is constructed to train the model to learn the consistent distribution of normal behavior across different modalities and domains. This training process is unsupervised, using only normal behavior samples. The core idea of contrastive learning is that for the same input (the normal behavior patch), its representations obtained through different views (time domain, structural domain) or different domains (time domain, frequency domain) should be similar or consistent with each other. Abnormal behavior, however, disrupts this consistency. The model's goal is to learn the boundaries of this normal consistency.
[0095] Loss function structure:
[0096]
[0097] For normal behavior, its time series pattern should be highly consistent with the underlying behavioral structure. For example, when a legitimate process starts, its API call sequence (timing) and the call relationships (structure) between APIs should exhibit a normal and expected pattern.
[0098]
[0099] The frequency characteristics of normal behavior should also maintain a high degree of similarity in their representations in the time and structural domains. For example, the periodic fluctuations in a normal process flow and the changes in structural complexity should exhibit a matching energy distribution in the frequency domain. This consistency will be broken if an attack introduces unexpected periodicity or frequency perturbations.
[0100]
[0101] Here, KL(·) represents the symmetric Kullback-Leibler divergence, an asymmetric measure of the difference between two probability distributions, which is made sensitive to both directions by symmetry; Ω(·) indicates that the stop-gradient operation is critical. This means that Ω(·) is treated as a constant when calculating the gradient. This is done to: prevent trivial solutions; and avoid the model directly applying Z... T Forced optimization to match Z G If they are completely identical, or vice versa, the two networks lose their independent feature extraction capabilities. Independent and complementary feature learning is encouraged: each branch (self-attention network and graph structure neural network) is encouraged to learn the most robust representation in its domain, while maintaining consistency with the representations of other branches. Stable training: This asymmetric gradient flow helps stabilize the training process of contrastive learning; d is the feature dimension; Let represent the j-th dimension frequency domain value; α∈[0,1] is a hyperparameter that can adjust the loss weight ratio. The model training phase uses only normal behavior samples. This means that pre-labeled attack samples are unnecessary, overcoming the problems of scarce attack samples and difficulty in obtaining labels in the cybersecurity field. The model constructs a robust representation of normal behavior patterns by learning the inherent consistency boundary of normal behavior.
[0102] In this embodiment of the application, in step S6, the inference stage processes the sample to be detected through a consistency detection model, calculates the consistency deviation between multimodal representations, and determines whether it is a fileless attack based on the consistency deviation and in conjunction with the anomaly detection judgment mechanism.
[0103] During the inference phase, when a new sequence of behaviors to be detected is input, it will sequentially undergo the aforementioned multimodal joint modeling and frequency domain transformation to obtain its time representation Z. TGraph structure representation of Z G and the corresponding frequency domain representation F T F G This invention uses the degree of inconsistency of behavior across different perspectives as an anomaly score. If the behavior to be detected deviates from the consistent distribution of the normal behavior pattern learned during the training phase, i.e., its multi-view representation shows significant differences in the feature space, its anomaly score will be high. Finally, this anomaly score is compared with a preset anomaly detection threshold; if it exceeds the threshold, it is determined to be a fileless attack behavior.
[0104] The anomaly detection threshold can be determined by statistical methods or cross-validation strategies on the training set, and can be set statically or dynamically and adaptively.
[0105] Specifically, during the inference phase, the input sequence to be detected is sequentially passed through a neural network with a self-attention mechanism to obtain a time-domain representation, and through a graph structure neural network to obtain a structural representation. The time-domain representation and the structural representation are then subjected to a Fast Fourier Transform to obtain the corresponding frequency-domain representation. A consistency metric between the time-domain representation and the structural representation in both the time and frequency domains is calculated, and the deviation of the consistency metric result is used as an anomaly score. If the anomaly score exceeds a set threshold, it is identified as a fileless attack.
[0106] The system receives a new sequence of behaviors to be detected. This sequence first undergoes the same data acquisition, unified encoding, patch partitioning, and behavior graph construction processes as in the training phase. Then, it sequentially passes through a self-attention neural network and a graph-structured neural network to obtain the representation Z. T With Z G And convert it to frequency domain representation F T With F G Next, we calculate the consistency deviation between the time domain and the frequency domain: Score = α·KL(Z T Z G )+(1-α)MAE(F T ,F G This anomaly score quantifies the degree to which the current behavior deviates from the learned consistency pattern of normal behavior in the time and frequency domains, representing a multi-view representation. A higher score indicates a greater deviation from the normal pattern. The calculated anomaly score is then compared to a pre-set threshold for decision-making. The anomaly detection threshold can be determined through statistical analysis or cross-validation strategies on an independent validation set (containing only normal samples). Furthermore, it can be statically set or dynamically adaptively adjusted based on the actual deployment environment and false positive / false negative tolerance. Finally, if the anomaly score exceeds the set threshold, the current behavior is determined to be a fileless attack, triggering corresponding alarms or defensive measures.
[0107] Through this multi-view, multi-domain joint modeling and contrastive learning mechanism, this invention can effectively capture subtle changes and complex correlations in the behavior of fileless attacks in memory, achieving efficient identification even against highly covert and adversarial threats. The fileless attack detection model based on multi-view behavior modeling and frequency domain enhanced contrastive learning is shown in the figure below. Figure 3 As shown.
[0108] Example 3, referring to Figure 4 This is one embodiment of the present invention, which provides a fileless attack detection system based on multi-view behavior modeling and frequency domain enhanced contrastive learning, including a data acquisition and encoding module, a behavior graph construction module, a feature extraction module, a frequency domain conversion module, a consistency modeling module, and an anomaly detection module.
[0109] The data acquisition and encoding module is used to collect and encode multi-source behavioral data of the target system during operation, and to perform unified encoding.
[0110] The behavior graph construction module is used to divide multi-source behavioral data into time series and construct corresponding behavior graphs.
[0111] The feature extraction module is used to input the segmented time series into a self-attention mechanism neural network to extract the temporal representation of the behavior, and input the constructed behavior graph into a graph structure neural network to extract the structural representation.
[0112] The frequency domain transformation module is used to perform fast Fourier transform on the time domain representation and the structural representation respectively to generate the frequency domain representation.
[0113] The consistency modeling module is used to construct a joint contrastive learning loss function, which includes a time domain loss that measures the difference between time domain representations using symmetric KL divergence, and a frequency domain loss that measures the difference between frequency domain representations using mean absolute error, to train the consistency detection model.
[0114] The anomaly detection module is used in the inference phase to process the sample to be detected through the consistency detection model, calculate the consistency deviation between multimodal representations, and determine whether it is a fileless attack based on the consistency deviation and the anomaly detection judgment mechanism.
[0115] This embodiment also provides an electronic device applicable to the fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning proposed in the above embodiment.
[0116] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, it implements the fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning as proposed in the above embodiments.
[0117] The storage medium proposed in this embodiment and the fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0118] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0119] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning, characterized in that: include, Collect and encode multi-source behavioral data of the target system during operation, and perform unified encoding; Multi-source behavioral data is divided into time series segments, and corresponding behavioral graphs are constructed. The segmented time series is input into a self-attention mechanism neural network to extract the temporal representation of the behavior, and the constructed behavior graph is input into a graph structure neural network to extract the structural representation. Fast Fourier transforms are performed on the time-domain representation and the structural representation respectively to generate the frequency-domain representation; Construct a joint contrastive learning loss function, which includes a time domain loss that measures the difference between time domain representations using symmetric KL divergence and a frequency domain loss that measures the difference between frequency domain representations using mean absolute error, and train a consistency detection model. During the inference phase, the sample to be tested is processed by the consistency detection model to calculate the consistency deviation between multimodal representations. Based on the consistency deviation and combined with the anomaly detection judgment mechanism, it is determined whether it is a fileless attack.
2. The fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning as described in claim 1, characterized in that: The self-attention mechanism neural network includes modeling the time domain representation and using a multi-head self-attention structure to extract time-dependent and sequence semantic features.
3. The fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning as described in claim 2, characterized in that: The graph structure neural network includes modeling the structural representation, extracting structural dependencies and contextual features between nodes by aggregating node adjacency information, and generating the structural representation.
4. The fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning as described in claim 3, characterized in that: The joint contrastive learning loss function includes a time-domain loss measured by symmetric KL divergence and a frequency-domain loss measured by mean absolute error. The process of training the consistency detection model is optimized using the joint contrastive learning loss function.
5. The fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning as described in claim 4, characterized in that: The time-domain representation refers to the representation generated after inputting the time series into a self-attention mechanism neural network. The self-attention mechanism neural network is a neural network based on a multi-head self-attention mechanism, which dynamically calculates the correlation between any two positions in the time series, extracts the temporal order, duration, and sequence expression features of behavioral events, and generates a representation vector that represents the temporal semantics of the time series.
6. The fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning as described in claim 4, characterized in that: The structural representation refers to the representation generated after the behavior graph is input into the graph structure neural network. The graph structure neural network performs feature aggregation and neighborhood modeling on the nodes in the behavior graph, learns the nonlinear structural dependencies between behavior events, and uses graph-level pooling to aggregate from all node representations to generate a structural representation vector representing the overall structure of the behavior graph.
7. The fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning as described in claim 4, characterized in that: The frequency domain representation refers to the representation generated by performing Fast Fourier Transform (FFT) on the time domain representation and the structural representation respectively. The FFT converts the time domain representation and the structural representation from time domain signals to frequency domain signals, obtaining the energy distribution at different frequencies, reflecting periodic fluctuations or frequency component changes. The frequency domain representation serves as the input to the frequency domain loss in the joint contrastive learning loss function.
8. A fileless attack detection system based on multi-view behavior modeling and frequency domain enhanced contrastive learning, employing the fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning as described in any one of claims 1 to 7, characterized in that, include: The system includes a data acquisition and encoding module, a behavior graph construction module, a feature extraction module, a frequency domain conversion module, a consistency modeling module, and an anomaly detection module. The data acquisition and encoding module is used to acquire and encode multi-source behavioral data of the target system during operation, and to perform unified encoding. The behavior graph construction module is used to divide multi-source behavior data into time series and construct corresponding behavior graphs; The feature extraction module is used to input the segmented time series into a self-attention mechanism neural network to extract the temporal representation of the behavior, and input the constructed behavior graph into a graph structure neural network to extract the structural representation. The frequency domain conversion module is used to perform fast Fourier transform on the time domain representation and the structural representation respectively to generate the frequency domain representation; The consistency modeling module is used to construct a joint contrastive learning loss function, which includes a time domain loss that measures the difference between time domain representations using symmetric KL divergence and a frequency domain loss that measures the difference between frequency domain representations using mean absolute error, to train the consistency detection model. The anomaly detection module is used in the inference phase to process the sample to be detected through the consistency detection model, calculate the consistency deviation between multimodal representations, and determine whether it is a fileless attack based on the consistency deviation and the anomaly detection judgment mechanism.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the fileless attack detection method based on multi-view behavior modeling and frequency domain enhanced contrastive learning as described in any one of claims 1 to 7.
Citation Information
Cited By
Serverless adaptive scheduling method and system based on consistency perception
CN121560499A
A serverless adaptive scheduling method and system based on consistency perception
CN121560499B