A multi-modal fusion graph attention-based malicious code homologous detection method and system

CN122508602BActive Publication Date: 2026-09-11NINGBO ZIHE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611006891.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-11
Estimated Expiration
2046-07-08

AI Technical Summary

Technical Problem

[0009]为了解决上述问题,本发明提供一种多模态融合图注意力的恶意代码同源检测方法及系统,解决现有技术中单一模态检测抗干扰能力弱、多模态浅层融合无法刻画模态间深层关联、静态图神经网络易受混淆技术干扰、加壳混淆场景下同源检测准确率显著下降的技术问题

Benefits of technology

本发明提出一种多模态融合图注意力的恶意代码同源检测方法、系统及设备,通过动态执行特征、静态控制流图与动态API调用序列的多模态信息协同,并借助API调用序列完成动态执行特征的时间窗口同步校准,既实现了不同维度特征的互补支撑,弥补了单一模态检测抗干扰能力不足的缺陷,也修正了动态采样的时间漂移问题,提升了动态特征的稳定性与样本间可比性;将两类特征映射至复数域并基于梯度场共轭关系生成跨模态调制场,能够显式刻画动态特征与静态结构间的相关性增强区域与抑制区域,突破了传统特征拼接、加权融合等浅层方式无法建模模态间结构性关联的局限,实现跨模态深层特征融合,有效增强了对代码混淆手段的抵抗能力;以跨模态调制场的空间梯度作为边权重调整项驱动图神经网络动态调整消息传播路径,打破了传统静态邻接聚合的固定传播模式,在静态控制流图因混淆发生失真时,可通过动态行为特征引导特征聚合方向,保障同源核心特征的有效提取;最终基于拓扑持续性特征与距离度量完成同源家族判定,依托拓扑特征的结构不变性适配恶意代码的局部变异扰动,显著提升了复杂混淆场景下恶意代码同源检测的准确率与鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122508602B_ABST
    Figure CN122508602B_ABST
Patent Text Reader

Abstract

The application discloses a multi-modal fusion graph attention malicious code homologous detection method and system, belonging to the network security technical field. Including obtaining the dynamic execution characteristics, static control flow graph and dynamic API calling sequence of the malicious code sample to be detected, using the API calling sequence to synchronize and calibrate the time window of the dynamic execution characteristics; mapping the two types of characteristics after calibration to the complex domain respectively, generating a cross-modal modulation field reflecting the correlation enhancement and inhibition area based on the conjugate relationship of the two gradient fields; inputting the spatial gradient of the cross-modal modulation field as the edge weight adjustment item into the graph neural network, dynamically adjusting the edge weight and iteratively updating the node embedding, obtaining the fused graph embedding; finally, extracting the topological persistence characteristics of the graph embedding, and determining the homologous family through the distance measurement with the known family reference topological characteristics; through the above scheme, the accuracy and robustness of homologous detection in the confusion scene are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of network security technology, and specifically discloses a method and system for detecting malicious code origins based on multimodal fusion graph attention. Background Technology

[0002] Malicious code homology detection is a core technology for tracing network threats, determining family affiliation, and analyzing attack chains. Its core objective is to determine whether there is a homology evolution relationship between unknown samples and known malicious families by mining the structural, behavioral, and semantic features of the code.

[0003] Existing malware same-origin detection technologies have the following main shortcomings:

[0004] First, many detection schemes rely on single-modal data, such as static detection based solely on static byte textures or control flow graphs, or dynamic detection based solely on dynamic API call sequences. These schemes cannot fully utilize the complementary information between different modalities. When samples are processed by adding packers or scrambling sections, single-modal features are prone to failure, resulting in a significant drop in detection performance.

[0005] Second, existing multimodal fusion schemes mostly adopt shallow fusion strategies such as feature splicing and weighted voting, which only perform simple combinations at the feature level or decision level. They do not perform in-depth modeling of the structural correlation and physical correspondence between modalities, and cannot characterize the enhancement and inhibition relationship between dynamic behavioral features and static structural features, resulting in very limited fusion gain.

[0006] Third, detection schemes based on graph neural networks generally use static adjacency matrices for message aggregation, and the node propagation path remains unchanged. This makes them susceptible to interference from techniques such as control flow flattening, false branch insertion, and instruction obfuscation. When the static graph structure becomes distorted due to obfuscation, the feature aggregation of the graph neural network will deviate, leading to a decrease in detection accuracy.

[0007] Fourth, in the context of tracing the origins of advanced persistent threats, malicious samples often undergo multiple layers of obfuscation and code mutation. Traditional methods of determination based on feature matching or embedding vector distance have low tolerance for topological perturbations, and the accuracy and recall of homology determination are difficult to meet the needs of practical scenarios.

[0008] Therefore, there is an urgent need for a malicious code same-origin detection scheme that can synergistically utilize three types of information—dynamic behavior, static structure, and temporal semantics—and possesses deep cross-modal fusion capabilities and anti-obfuscation robustness. Summary of the Invention

[0009] To address the aforementioned issues, this invention provides a method and system for detecting malicious code origins using multimodal fusion graph attention. This addresses the technical problems in existing technologies, such as weak anti-interference capability of single-modal detection, inability of shallow multimodal fusion to characterize deep intermodal relationships, susceptibility of static graph neural networks to obfuscation techniques, and significant decrease in the accuracy of origin detection under obfuscation scenarios.

[0010] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for detecting the same source of malicious code using multimodal fusion graph attention, the method comprising: The dynamic execution features, static control flow graph, and dynamic API call sequence of the malicious code sample to be detected are obtained. The time window of the dynamic execution features is synchronously calibrated using the API call sequence. The node features of the static control flow graph structure are embedded into hyperbolic space through exponential mapping to obtain the hyperbolic embedding representation of the static control flow graph structure. The synchronized calibrated dynamic execution features and the static control flow graph are mapped to the complex domain, respectively. A cross-modal modulation field is generated based on the conjugate relationship of their gradient fields in the complex domain. The cross-modal modulation field reflects the correlation enhancement region and suppression region between the dynamic features and the static structure. The spatial gradient of the cross-modal modulation field is used as an edge weight adjustment term and input into the graph neural network. The graph neural network uses the static control flow graph as the initial topology and dynamically adjusts the edge weights according to the edge weight adjustment term during message passing, iteratively updating the node embedding to obtain the fused graph embedding. Extract the topological persistence features of the graph embedding, calculate the distance metric between the topological persistence features and the known family reference topological features, and determine the homologous family based on the comparison result of the distance metric and a preset threshold.

[0011] Optionally, obtaining the dynamic execution characteristics of the malicious code sample to be detected includes: The sample to be tested is dynamically executed in a sandbox environment. Memory page data is captured in fixed time windows, and the information entropy of each memory page is calculated. The entropy value sequence of multiple consecutive time windows is arranged in memory address order to generate a three-dimensional entropy tensor as the dynamic execution feature.

[0012] Optionally, obtaining the static control flow graph of the malicious code sample to be detected includes: Obtain the static control flow graph and function call graph of the sample, extract the node features of the static control flow graph and the function call graph, calculate the spectral entropy of the graph Laplacian matrix, embed the node features into the hyperbolic space through exponential mapping, take negative curvature, and obtain the hyperbolic embedding representation of the static control flow graph. The hyperbolic embedding representation is used in subsequent steps to be jointly mapped to the complex domain with the synchronously calibrated dynamic execution features.

[0013] Optionally, the step of synchronously calibrating the time window of the dynamic execution feature using the API call sequence includes: The API call types in the dynamic API call sequence are mapped to a vector sequence through an embedding layer. The vector sequence is input into a neural differential equation solver composed of a fully connected network. An adaptive step size is used to solve for the time offset calibration amount. The time window step size of the dynamic execution feature is corrected according to the time offset calibration amount. The memory page entropy value sequence is re-acquired to update the dynamic execution feature. The updated dynamic execution feature is used as the synchronously calibrated dynamic execution feature.

[0014] Optionally, mapping the synchronized calibrated dynamic execution features and the static control flow graph to the complex domain includes: The synchronized and calibrated dynamic execution features are mapped to the complex domain to obtain the first complex domain features; the hyperbolic embedding representation of the static control flow graph is mapped to the complex domain to obtain the second complex domain features. Both the first complex field feature and the second complex field feature are represented as a combination of real and imaginary parts.

[0015] Optionally, generating the cross-modal modulation field based on the conjugate relationship of the gradient fields of the two in the complex domain includes: Automatic differentiation is performed on the first complex domain feature and the second complex domain feature respectively to obtain their respective gradient fields. Alignment interpolation is performed on the gradient fields on the same sampling grid. The conjugate product of the gradient fields of the first complex domain feature and the second complex domain feature is calculated. The imaginary part of the conjugate product is taken as the cross-mode modulation field.

[0016] Optionally, the step of dynamically adjusting the edge weights according to the edge weight adjustment term and iteratively updating the node embedding during message passing includes: For any node in the current graph topology, the attention coefficient between the node and its neighboring nodes is calculated. The spatial gradient of the cross-modal modulation field is mapped to the edge weight adjustment value. The edge weight adjustment value is superimposed as a bias term onto the linear transformation result of the neighboring node features. The neighboring node features of the node are weighted and aggregated using the superimposed attention coefficient. After transformation by the activation function, the updated embedding vector of the node is obtained. The updated embedding vector is iteratively updated to obtain the fused graph embedding.

[0017] Optionally, mapping the synchronized calibrated dynamic execution features and the static control flow graph to the complex domain further includes: After obtaining the first complex domain feature and the second complex domain feature, an orthogonal constraint is applied to the first complex domain feature and the second complex domain feature. The orthogonal constraint is as follows: the real part features corresponding to the first complex domain feature and the second complex domain feature are orthogonalized in the real part space, and the orthogonalization loss is used as an additional regularization term for training the graph neural network. The orthogonal constraint imposes a feature independence restriction between the complex domain mapping and the cross-modal modulation field generation.

[0018] Optionally, extracting the topological persistence features of the graph embedding includes: The fused graph embedding is mapped back to Euclidean space. A filtering threshold is set using the distance between the graph embedding points in the Euclidean space as a metric. Below the filtering threshold, all point pairs with a distance less than the filtering threshold are connected to form a simple complex. The persistence barcode with 0-dimensional and 1-dimensional Betti numbers is extracted. The persistence barcode is converted into topological persistence features in the form of discrete point sets.

[0019] Secondly, this invention provides a multimodal fusion graph attention-based malicious code homology detection system, comprising: The synchronous calibration module is used to acquire the dynamic execution features, static control flow graph, and dynamic API call sequence of the malicious code sample to be detected, and to synchronously calibrate the time window of the dynamic execution features using the API call sequence; the node features of the static control flow graph structure are embedded into hyperbolic space through exponential mapping to obtain the hyperbolic embedding representation of the static control flow graph structure. The generation module is used to map the synchronously calibrated dynamic execution features and the static control flow graph to the complex domain respectively, and generate a cross-modal modulation field based on the conjugate relationship of the gradient fields of the two in the complex domain. The cross-modal modulation field reflects the correlation enhancement region and suppression region between the dynamic features and the static structure. The iterative update module is used to input the spatial gradient of the cross-modal modulation field as an edge weight adjustment term into the graph neural network. The graph neural network uses the static control flow graph as the initial topology and dynamically adjusts the edge weights according to the edge weight adjustment term during message passing, iteratively updating the node embedding to obtain the fused graph embedding. The determination calculation module is used to extract the topological persistence features of the graph embedding, calculate the distance metric between the topological persistence features and the known family reference topological features, and determine the homologous family based on the comparison result of the distance metric and the preset threshold.

[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention proposes a multimodal fusion graph attention-based method, system, and device for malware homology detection. By coordinating multimodal information from dynamic execution features, static control flow graphs, and dynamic API call sequences, and using API call sequences to synchronize and calibrate the time window of dynamic execution features, it achieves complementary support from different feature dimensions, compensating for the insufficient anti-interference capability of single-modal detection, and correcting the time drift problem of dynamic sampling, thus improving the stability of dynamic features and the comparability between samples. Mapping two types of features to the complex domain and generating a cross-modal modulation field based on the gradient field conjugate relationship can explicitly characterize the correlation enhancement and suppression regions between dynamic features and static structures, breaking through the limitations of traditional feature splicing and addition... Shallow methods such as weighted fusion cannot model the structural relationships between modalities. This paper achieves deep feature fusion across modalities, effectively enhancing resistance to code obfuscation techniques. By using the spatial gradient of the cross-modal modulation field as the edge weight adjustment term to drive the graph neural network to dynamically adjust the message propagation path, it breaks the fixed propagation mode of traditional static adjacency aggregation. When the static control flow graph is distorted due to obfuscation, it can guide the feature aggregation direction through dynamic behavioral features, ensuring the effective extraction of core features of the same origin. Finally, it completes the determination of the same origin family based on topological persistence features and distance metrics. Relying on the structural invariance of topological features, it adapts to the local mutation perturbation of malicious code, significantly improving the accuracy and robustness of malicious code same origin detection in complex obfuscation scenarios. Attached Figure Description

[0021] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0022] Figure 1 This is a flowchart of a malicious code homology detection method based on multimodal fusion graph attention provided in an embodiment of the present invention; Figure 2This is a flowchart of calculating the cross-modal modulation field through the conjugate product of gradient fields provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of a malicious code homology detection system based on multimodal fusion graph attention provided in an embodiment of the present invention; Figure 4 This is an internal structure diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0023] The embodiments of the technical solution of the present invention will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of the present invention and are therefore merely examples, and should not be construed as limiting the scope of protection of the present invention.

[0024] It should be noted that, unless otherwise stated, the technical or scientific terms used in this invention should have the ordinary meaning as understood by one of ordinary skill in the art.

[0025] This invention provides a method and system for detecting the origin of malicious code using multimodal fusion graph attention, applicable to applications of tracing the origins of malicious code families. The embodiments of this invention are described below with reference to the accompanying drawings.

[0026] Example 1: As Figure 1 As shown, one embodiment of the present invention provides a method for detecting the same source of malicious code using multimodal fusion graph attention. This method specifically includes the following steps: S101 acquires the dynamic execution features, static control flow graph, and dynamic API call sequence of the malicious code sample to be detected, and uses the API call sequence to synchronously calibrate the time window of the dynamic execution features; embeds the node features of the static control flow graph structure into hyperbolic space through exponential mapping to obtain the hyperbolic embedding representation of the static control flow graph structure; S102 maps the synchronously calibrated dynamic execution features and the static control flow graph to the complex domain respectively, and generates a cross-modal modulation field based on the conjugate relationship of their gradient fields in the complex domain. The cross-modal modulation field reflects the correlation enhancement region and suppression region between the dynamic features and the static structure. S103 uses the spatial gradient of the cross-modal modulation field as an edge weight adjustment term and inputs it into the graph neural network. The graph neural network uses the static control flow graph as the initial topology and dynamically adjusts the edge weights according to the edge weight adjustment term during message passing, iteratively updates the node embedding, and obtains the fused graph embedding. S104 Extracts the topological persistence features of the graph embedding, calculates the distance metric between the topological persistence features and the known family reference topological features, and determines the homologous family based on the comparison result of the distance metric and a preset threshold.

[0027] In step S101 above, obtaining the dynamic execution characteristics of the malicious code sample to be detected includes: The sample to be tested is dynamically executed in a sandbox environment. Memory page data is captured in fixed time windows, and the information entropy of each memory page is calculated. The entropy value sequence of multiple consecutive time windows is arranged in memory address order to generate a three-dimensional entropy tensor as the dynamic execution feature.

[0028] In one embodiment, dynamic execution feature extraction selects a malicious code sample in PE format to be detected, loads it into a virtual execution environment built on a Cuckoo sandbox, sets the dynamic execution duration to 120 seconds, and synchronously collects a sequence of memory page snapshots with an initial time window step of Δt0 = 5 seconds. The memory pages within each time window are divided into 4KB blocks, and the information entropy of each memory block is calculated using the following formula: ;in, Let be the frequency of the i-th byte value in the memory page. This is used to measure the amount of information carried when a specific byte value appears; the entropy value sequence of 24 consecutive time windows is arranged in memory address order to generate a three-dimensional entropy tensor of size H=64, W=64, C=24, which serves as the dynamic execution feature. This tensor simultaneously encodes the spatial arrangement and temporal dissipation characteristics of code execution.

[0029] In step S101 above, obtaining the static control flow graph of the malicious code sample to be detected includes: Obtain the static control flow graph and function call graph of the sample, extract the node features of the static control flow graph and the function call graph, calculate the spectral entropy of the graph Laplacian matrix, embed the node features into the hyperbolic space through exponential mapping, take negative curvature, and obtain the hyperbolic embedding representation of the static control flow graph. The hyperbolic embedding representation is used in subsequent steps to be jointly mapped to the complex domain with the synchronously calibrated dynamic execution features.

[0030] In one embodiment, IDA Pro or Ghidra disassemblers are used to perform static analysis on the sample to be tested, extracting the control flow graph and function call graph of the sample. The Laplacian matrix of the computational graph is L=D. A, where D is the degree matrix and A is the adjacency matrix, calculates the spectral entropy of the graph based on the eigenvalues ​​of the Laplacian matrix, and uses it as the initial structural feature of the nodes.

[0031] The initial features of the nodes are embedded into a Poincaré hyperbolic space using an exponential mapping, with the hyperbolic space curvature set to c=-1 and the embedding dimension set to 512, to obtain a hierarchical potential tensor, which serves as the hyperbolic embedding representation of the static control flow graph. Hyperbolic space embedding can naturally preserve the hierarchical nesting relationships of the graph structure and is more suitable for the tree-and-network hybrid structure of malicious code function calls.

[0032] Step S101 above, obtaining the dynamic API call sequence of the malicious code sample to be detected includes: The sample to be tested is dynamically executed in a sandbox environment. The application programming interfaces called by the sample during the process and their call timestamps are captured by API hooking or system call monitoring. The type and parameters of each API call are recorded, and an API call log is generated in chronological order. The dynamic API call sequence is extracted from the API call log.

[0033] In step S101 above, synchronizing the time window of the dynamic execution feature using the API call sequence includes: The API call types in the dynamic API call sequence are mapped to a vector sequence through an embedding layer. The vector sequence is input into a neural differential equation solver composed of a fully connected network. An adaptive step size is used to solve for the time offset calibration amount. The time window step size of the dynamic execution feature is corrected according to the time offset calibration amount. The memory page entropy value sequence is re-acquired to update the dynamic execution feature. The updated dynamic execution feature is used as the synchronously calibrated dynamic execution feature.

[0034] In one embodiment, a dynamic API call sequence is extracted from the API hook logs of the sandbox, and the API call type is mapped to a 128-dimensional vector sequence through an embedding layer. This vector sequence is then input into a neural differential equation solver consisting of two fully connected layers, and the continuous trajectory flow is solved using the Dormand-Prince method with an adaptive step size to calculate the time offset calibration amount Δt.

[0035] If Δt > 0.5 seconds, the time window step of the dynamic execution feature is corrected to Δt0 + Δt, where Δt is the time offset calibration amount. The memory page entropy acquisition and 3D entropy tensor generation steps are then re-executed to obtain the synchronously calibrated dynamic execution feature. This calibration can correct the time window drift caused by differences in execution speed among different samples, improving the consistency of dynamic features across different samples.

[0036] In step S102 above, mapping the synchronized calibrated dynamic execution features and the static control flow graph to the complex domain includes: The synchronized and calibrated dynamic execution features are mapped to the complex domain to obtain the first complex domain features; the hyperbolic embedding representation of the static control flow graph is mapped to the complex domain to obtain the second complex domain features. Both the first complex field feature and the second complex field feature are represented as a combination of real and imaginary parts.

[0037] In the above embodiments, a complex encoder with a dual-channel structure of real and imaginary parts is used to perform complex domain mapping on the hyperbolic embedding representation of the synchronously calibrated dynamic execution features and the static control flow graph respectively: For the dynamic execution features, multi-scale spatiotemporal texture features are extracted through a first 3D convolutional sub-network, and then real and imaginary tensors are generated through a complex mapping layer to obtain the first complex domain features. For the hyperbolic embedding representation of static control flow graphs, topological features are extracted through a graph attention subnetwork, and then real and imaginary tensors are generated through a complex mapping layer to obtain the second complex domain features. ; in, The amplitude of the static mode. For the amplitude of the dynamic mode, For dynamic phase, It is a static phase.

[0038] In this embodiment, the complex space dimension d=256, where both the real and imaginary parts are 128-dimensional. After complex encoding, the features simultaneously contain amplitude and phase information, providing a mathematical basis for subsequent modeling of intermodal interference relationships.

[0039] In the above embodiment S102, the step of mapping the synchronized calibrated dynamic execution features and the static control flow graph to the complex domain respectively further includes: After obtaining the first complex domain feature and the second complex domain feature, an orthogonal constraint is applied to the first complex domain feature and the second complex domain feature. The orthogonal constraint is as follows: the real part features corresponding to the first complex domain feature and the second complex domain feature are orthogonalized in the real part space, and the orthogonalization loss is used as an additional regularization term for training the graph neural network. The orthogonal constraint imposes a feature independence restriction between the complex domain mapping and the cross-modal modulation field generation.

[0040] In the above embodiment S102, generating a cross-mode modulation field based on the conjugate relationship of the gradient fields of the two in the complex domain includes: Automatic differentiation is performed on the first complex domain feature and the second complex domain feature respectively to obtain their respective gradient fields. Alignment interpolation is performed on the gradient fields on the same sampling grid. The conjugate product of the gradient fields of the first complex domain feature and the second complex domain feature is calculated. The imaginary part of the conjugate product is taken as the cross-mode modulation field.

[0041] In one embodiment, after generating the first complex domain features and the second complex domain features, orthogonalization is applied to the two types of features in the real part space, and the corresponding orthogonal loss function is... for:

[0042] in, Let I be the Frobenius norm, and I be the identity matrix; The transpose of the matrix; Second complex field characteristics; The orthogonal loss serves as an additional regularization term during model training, constraining the independence of features between two modalities, avoiding information redundancy between modalities, and improving the effectiveness of fusion.

[0043] In one embodiment, such as Figure 2 As shown, based on the structural characteristics of the steady-state solution of the Helmholtz equation, the cross-mode modulation field is calculated through the conjugate product of the gradient field. The specific steps are as follows: S21 respectively applies the feature Z of the first complex field v Second complex field feature Z g Perform automatic differentiation to obtain their respective gradient fields. Z v and Z g ; S22 performs aligned interpolation of the two gradient fields on the same sampling grid to ensure a one-to-one correspondence in spatial dimensions; S23 calculates the conjugate product of the characteristic gradient fields of the first complex domain and the second complex domain, and takes the imaginary part of the result as the cross-modal modulation field. The calculation formula is as follows: ; Where ⊙ represents the Hadamarda accumulation. for The conjugate of , where Im represents taking the imaginary part.

[0044] The cross-modal modulation field includes interference enhancement and interference cancellation regions, reflecting the local correlation between dynamic execution features and static graph structure features: the enhancement regions correspond to structural locations where the two types of features are highly matched, while the cancellation regions correspond to locations where the static structure may exhibit confusion or distortion. In this embodiment, the fixed-point iterative method is used to solve the approximate Helmholtz equation, with an initial iteration step count K=5 and a residual convergence threshold of 1×10⁻⁶. 4 The computation is performed using the GPU-based cuBLAS library to perform batch complex matrix multiplications to improve computational efficiency.

[0045] In the above embodiment S103, the step of dynamically adjusting the edge weights according to the edge weight adjustment item and iteratively updating the node embedding during message transmission includes: For any node in the current graph topology, the attention coefficient between the node and its neighboring nodes is calculated. The spatial gradient of the cross-modal modulation field is mapped to the edge weight adjustment value. The edge weight adjustment value is superimposed as a bias term onto the linear transformation result of the neighboring node features. The neighboring node features of the node are weighted and aggregated using the superimposed attention coefficient. After transformation by the activation function, the updated embedding vector of the node is obtained. The updated embedding vector is iteratively updated to obtain the fused graph embedding.

[0046] Based on the message propagation path adjustment graph neural network of cross-modal modulation field, a graph embedding that integrates multimodal information is generated.

[0047] Using a static control flow graph as the initial topology, the spatial gradient of the cross-modal modulation field is injected into the message passing process of the graph attention network as an edge weight adjustment term. For any node i in the graph, its embedding update formula at layer l+1 is: ; Where: N(i) is the set of neighboring nodes of node i; represents the original attention coefficients between node i and node j; W is the trainable feature transformation matrix. β is the gradient of the cross-modal modulation field to the spatial location of node i, i.e., the edge weight adjustment value; β is the trainable coupling coefficient, initialized to 0.1; σ is the nonlinear activation function.

[0048] In this embodiment, the graph attention network is set to 3 layers, with a hidden layer dimension of 256. After 3 rounds of message passing iterations, the updated graph embedding H is output. 3 ∈ n×256 , where N is the number of nodes in the graph. This graph embedding simultaneously integrates the topological information of the static structure with the behavioral information of dynamic execution, and the propagation path is guided by dynamic features, exhibiting strong robustness against static structure obfuscation.

[0049] In the above embodiment S104, the extraction of the topological persistence features of the graph embedding includes: mapping the fused graph embedding back to Euclidean space, setting a filtering threshold using the distance between graph embedding points in the Euclidean space as a metric, connecting all point pairs with a distance less than the filtering threshold below the filtering threshold to form a simple complex, extracting the persistence barcode of 0-dimensional and 1-dimensional Betti numbers, and converting the persistence barcode into topological persistence features in the form of discrete point sets.

[0050] In the above embodiment S104, calculating the distance metric between the topological persistence feature and the known family reference topological feature includes: The topological persistence features of the sample to be detected and the known family reference topological features are discretized into finite point sets, respectively. The 2-Wasserstein distance between the finite point sets is calculated using the Sinkhorn approximation algorithm as the distance metric. The distance metric is used to compare with the preset threshold in subsequent steps.

[0051] Further, the step of determining the homologous family based on the comparison result of the distance metric and the preset threshold includes: The distance metric is compared with a preset homology threshold. When the distance metric is less than the preset homology threshold, the sample to be detected is determined to be homologous with the corresponding known family, and the family identifier and confidence level are output.

[0052] In one embodiment, the fused graph embedding is mapped from hyperbolic space back to Euclidean space, and a Vietoris-Rips simple complex is constructed based on the node embedding points in Euclidean space: The distance metric used is Euclidean distance; The filtering threshold is adaptively set based on the 20th percentile of the distance between embedding points; The maximum simplex dimension is set to 1 to balance computational overhead and feature representation capability.

[0053] The persistent barcodes with 0-dimensional and 1-dimensional Betty numbers are extracted, and the bar intervals are converted into discrete point sets to obtain the topological persistence features. These topological persistence features can characterize the overall topological structure of the point set, are invariant to perturbations, additions, and deletions of local nodes, and adapt to structural changes after malicious code mutations.

[0054] Reference topological persistence features are pre-constructed for each known malware family. For the topological persistence features of the sample to be detected, the topological persistence features of the sample to be detected are discretized into a first point set, and the reference topological features of each known family are discretized into a second point set. The 2-Wasserstein distance between the first point set and the second point set is calculated. This embodiment uses the Sinkhorn iterative approximation method, with the regularization coefficient set to 0.01.

[0055] Select the calculated minimum Wasserstein distance d min Compare it with the preset homology threshold τ: if d min If the value is less than τ, then the sample to be detected is determined to be homologous with the corresponding family, the corresponding family identifier is output, and the confidence score is calculated using the following formula: p=softmax(-τ) min The preset homology threshold τ is determined in advance by maximizing the F1-score on the validation set.

[0056] In a preferred embodiment, to further improve the reliability of the determination result, the present invention also introduces a feedback iteration mechanism: After obtaining the homology determination result and its confidence level p, if the determination confidence level p < 0.85, and the number of iterations K for the current cross-modal modulation field has not reached the preset maximum upper limit K. max (For example, K) max If the number of iterations reaches 5, a feedback signal is sent, the iteration number K = K + 1, and the process returns to perform the regeneration step of the cross-modal modulation field and its subsequent steps until the confidence threshold is met or the maximum iteration limit is reached. If the confidence requirement is not met even after the maximum number of iterations is reached, the current best family result and low confidence flag are output. This mechanism further ensures the robustness of detection in complex and confusing scenarios through dynamic iterative optimization.

[0057] Example 2: Based on the same technical concept, Example 2 of this invention also provides a multimodal fusion graph attention-based malicious code homology detection system, such as... Figure 3 As shown, it includes: a synchronous calibration module 210, a generation module 220, an iterative update module 230, and a judgment calculation module 240, wherein: The synchronous calibration module 210 is used to acquire the dynamic execution features, static control flow graph, and dynamic API call sequence of the malicious code sample to be detected, and to perform synchronous calibration on the time window of the dynamic execution features using the API call sequence; and to embed the node features of the static control flow graph structure into hyperbolic space through exponential mapping to obtain the hyperbolic embedding representation of the static control flow graph structure. The generation module 220 is used to map the synchronously calibrated dynamic execution features and the static control flow graph to the complex domain respectively, and generate a cross-modal modulation field based on the conjugate relationship of the gradient fields of the two in the complex domain. The cross-modal modulation field reflects the correlation enhancement region and suppression region between the dynamic features and the static structure. The iterative update module 230 is used to input the spatial gradient of the cross-modal modulation field as an edge weight adjustment term into the graph neural network. The graph neural network uses the static control flow graph as the initial topology and dynamically adjusts the edge weights according to the edge weight adjustment term during message passing, iteratively updates the node embedding, and obtains the fused graph embedding. The determination calculation module 240 is used to extract the topological persistence features of the graph embedding, calculate the distance metric between the topological persistence features and the known family reference topological features, and determine the homologous family based on the comparison result of the distance metric and the preset threshold.

[0058] Example 3: In one embodiment, Example 3 of the present invention also provides an electronic device; the electronic device may be a terminal, and its internal structure diagram may be as follows. Figure 4As shown. The electronic device includes a processor, memory, communication interface, display screen, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements the malicious code homology detection method based on multimodal fusion graph attention as described in any one of steps S101 to S104. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input device can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.

[0059] Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0060] Embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0061] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0062] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0063] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0064] The above are merely embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall fall within the protection scope of the present invention.

Claims

1. A method for detecting the same source of malicious code using multimodal fusion graph attention, characterized in that, Includes the following steps: Obtain the dynamic execution characteristics, static control flow graph, and dynamic API call sequence of the malicious code sample to be detected; The API call types in the dynamic API call sequence are mapped to vector sequences through an embedding layer. The vector sequences are input into a neural differential equation solver composed of a fully connected network. An adaptive step size is used to solve for the time offset calibration amount. The time window step size of the dynamic execution feature is corrected according to the time offset calibration amount. The memory page entropy sequence is re-acquired to update the dynamic execution feature. The updated dynamic execution feature is used as the synchronously calibrated dynamic execution feature. The node features of the static control flow graph structure are embedded into hyperbolic space through exponential mapping to obtain the hyperbolic embedding representation of the static control flow graph structure. The synchronized and calibrated dynamic execution features are mapped to the complex domain to obtain the first complex domain features; The hyperbolic embedding representation of the static control flow graph is mapped to the complex domain to obtain a second complex domain feature. Automatic differentiation is performed on the first and second complex domain features to obtain their respective gradient fields. Alignment interpolation is performed on the gradient fields on the same sampling grid. The conjugate product of the gradient fields of the first and second complex domain features is calculated. The imaginary part of the conjugate product is taken as the cross-modal modulation field, which reflects the correlation enhancement and suppression regions between dynamic features and static structure. The spatial gradient of the cross-modal modulation field is used as an edge weight adjustment term and input into the graph neural network. The graph neural network uses the static control flow graph as the initial topology and dynamically adjusts the edge weights according to the edge weight adjustment term during message passing, iteratively updating the node embedding to obtain the fused graph embedding. Extract the topological persistence features of the graph embedding, calculate the distance metric between the topological persistence features and the known family reference topological features, and determine the homologous family based on the comparison result of the distance metric and a preset threshold.

2. The method according to claim 1, characterized in that, The dynamic execution characteristics of the malicious code sample to be detected include: The sample to be tested is dynamically executed in a sandbox environment. Memory page data is captured in fixed time windows, and the information entropy of each memory page is calculated. The entropy value sequence of multiple consecutive time windows is arranged in memory address order to generate a three-dimensional entropy tensor as the dynamic execution feature.

3. The method according to claim 1, characterized in that, The static control flow graph for obtaining the malicious code sample to be detected includes: Obtain the static control flow graph and function call graph of the sample, extract the node features of the static control flow graph and the function call graph, calculate the spectral entropy of the graph Laplacian matrix, embed the node features into the hyperbolic space through exponential mapping, take negative curvature, and obtain the hyperbolic embedding representation of the static control flow graph. The hyperbolic embedding representation is used in subsequent steps to be jointly mapped to the complex domain with the synchronously calibrated dynamic execution features.

4. The method according to claim 1, characterized in that, Both the first complex field feature and the second complex field feature are represented as a combination of real and imaginary parts.

5. The method according to claim 1, characterized in that, The step of dynamically adjusting edge weights based on the edge weight adjustment item and iteratively updating node embeddings during message passing includes: For any node in the current graph topology, the attention coefficient between the node and its neighboring nodes is calculated. The spatial gradient of the cross-modal modulation field is mapped to the edge weight adjustment value. The edge weight adjustment value is superimposed as a bias term onto the linear transformation result of the neighboring node features. The neighboring node features of the node are weighted and aggregated using the superimposed attention coefficient. After transformation by the activation function, the updated embedding vector of the node is obtained. The updated embedding vector is iteratively updated to obtain the fused graph embedding.

6. The method according to claim 1, characterized in that, The step of mapping the synchronized calibrated dynamic execution features and the static control flow graph to the complex domain also includes: After obtaining the first complex domain feature and the second complex domain feature, an orthogonal constraint is applied to the first complex domain feature and the second complex domain feature. The orthogonal constraint is as follows: the real part features corresponding to the first complex domain feature and the second complex domain feature are orthogonalized in the real part space, and the orthogonalization loss is used as an additional regularization term for training the graph neural network. The orthogonal constraint imposes a feature independence restriction between the complex domain mapping and the cross-modal modulation field generation.

7. The method according to claim 1, characterized in that, The extraction of the topological persistence features of the graph embedding includes: mapping the fused graph embedding back to Euclidean space; setting a filtering threshold using the distance between graph embedding points in the Euclidean space as a metric; connecting all point pairs with a distance less than the filtering threshold below the filtering threshold to form a simple complex; extracting the persistence barcode of 0-dimensional and 1-dimensional Betti numbers; and converting the persistence barcode into topological persistence features in the form of discrete point sets.

8. A malicious code homology detection system based on multimodal fusion graph attention, characterized in that, include: The synchronous calibration module is used to acquire the dynamic execution characteristics, static control flow graph, and dynamic API call sequence of the malicious code sample to be detected. The API call types in the dynamic API call sequence are mapped to a vector sequence through an embedding layer. The vector sequence is then input into a neural differential equation solver composed of a fully connected network. An adaptive step size is used to obtain the time offset calibration. The time window step size of the dynamic execution feature is corrected based on the time offset calibration. The memory page entropy sequence is re-acquired to update the dynamic execution feature. The updated dynamic execution feature is used as the synchronously calibrated dynamic execution feature. The node features of the static control flow graph structure are embedded into hyperbolic space through exponential mapping to obtain a hyperbolic embedding representation of the static control flow graph structure. The generation module is used to map the synchronized and calibrated dynamic execution features to the complex domain to obtain the first complex domain features; The hyperbolic embedding representation of the static control flow graph is mapped to the complex domain to obtain a second complex domain feature. Automatic differentiation is performed on the first and second complex domain features to obtain their respective gradient fields. Alignment interpolation is performed on the gradient fields on the same sampling grid. The conjugate product of the gradient fields of the first and second complex domain features is calculated. The imaginary part of the conjugate product is taken as the cross-modal modulation field, which reflects the correlation enhancement and suppression regions between dynamic features and static structure. The iterative update module is used to input the spatial gradient of the cross-modal modulation field as an edge weight adjustment term into the graph neural network. The graph neural network uses the static control flow graph as the initial topology and dynamically adjusts the edge weights according to the edge weight adjustment term during message passing, iteratively updating the node embedding to obtain the fused graph embedding. The determination calculation module is used to extract the topological persistence features of the graph embedding, calculate the distance metric between the topological persistence features and the known family reference topological features, and determine the homologous family based on the comparison result of the distance metric and the preset threshold.

Citation Information

Patent Citations

  • Adaptive fuzzy test method and system based on deep learning

    CN121808798A

  • Software development application data processing method and system based on artificial intelligence

    CN121833039A