A malicious code variant detection method, system and device for controlling semantic matching of a control flow graph

CN122333470BActive Publication Date: 2026-09-18NINGBO ZIHE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610812680.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-08
Publication Date
2026-09-18
Estimated Expiration
2046-06-08

AI Technical Summary

Technical Problem

一方面,仅依赖基本块局部语义向量与压缩后拓扑一致性进行收敛,仍可能受到桥接块注入、语义稀释块复制、调度器块迁移等策略影响,使部分本应稳定识别的核心子结构在不同变种之间出现规范排序漂移

Benefits of technology

本申请提出一种控制流图语义匹配的恶意代码变种检测方法、系统及设备,通过引入主锚点与主干边,构建跨变形的稳定骨架,从根源上抵御桥接块、调度块及伪锚点等干扰因素。同时,采用锚点约束压缩与噪声延迟吸收机制,有效避免语义污染,确保语义核心图的稳定收敛。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122333470B_ABST
    Figure CN122333470B_ABST
Patent Text Reader

Abstract

The application discloses a malicious code variant detection method and system based on control flow graph semantic matching, and a device thereof, and belongs to the technical field of computer network security. The method comprises the following steps: disassembling a target executable file, constructing a control flow graph, extracting basic block semantic features to generate a semantic vector, clustering based on the semantic vector to obtain a semantic equivalence class, identifying a main anchor point and constructing a main anchor point skeleton graph, performing semantic compression and noise processing with the anchor point skeleton as a constraint, obtaining a semantic core graph through topological verification, layering and normalizing the semantic core graph and the main anchor point skeleton graph to generate a composite hash signature, and matching the composite hash signature with a malicious code family signature library to complete variant determination. The above scheme can resist structure confusion disturbance and improve the consistency of homologous variant recognition through the main anchor point constraint compression and normalization process, and has the advantages of strong robustness, high efficiency, good interpretability and the like, and is suitable for malicious code variant detection based on control flow graph semantic matching in a complex confusion environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer network security technology, and specifically discloses a method, system and device for detecting malicious code variants by control flow graph semantic matching. Background Technology

[0002] Malicious code variants typically alter the program's apparent structure through methods such as packing, instruction equivalence substitution, basic block splitting and merging, control flow flattening, dead code insertion, call bridging, and scheduling block rewriting to evade traditional detection methods based on signatures, graph edit distance, graph kernel, or original control flow graph matching. Especially in complex obfuscated scenarios, attackers often do not change the core malicious behavior, but rather create structural noise, local semantic drift, and graph representation instability, causing variants within the same family to exhibit significant topological differences at the control flow graph level.

[0003] While existing solutions have proposed compressing the control flow graph into a semantic core graph based on the semantic equivalence of basic blocks and generating immutable signatures based on normalized hashing, there is still room for further optimization under higher-intensity adversarial examples. On the one hand, relying solely on the consistency of the local semantic vectors of basic blocks with the compressed topology for convergence can still be affected by strategies such as bridging block injection, semantic dilution block duplication, and scheduler block migration, causing some core substructures that should be stably identified to experience normalized order drift between different variants. On the other hand, although a single final-state signature facilitates fast matching, it lacks explicit constraints on the stability at different stages of the compression process, leading to the possibility of signature splitting of homologous samples or mis-aggregation of closely related but non-homologous samples among samples across compilers, optimization levels, and shell peeling depths. Summary of the Invention

[0004] To address the aforementioned issues, this application provides a method, system, and device for detecting malicious code variants by controlling flow graph semantic matching. Through main anchor point identification, backbone edge construction, anchor point constraint compression, noise buffering, hierarchical normalization, and compound signature generation, stable detection under high-intensity obfuscation is achieved.

[0005] To achieve the above-mentioned objectives, this application adopts the following technical solution: Firstly, this application provides a method for detecting malicious code variants by controlling flow graph semantic matching, including: The target executable file is disassembled, a control flow graph is constructed, and the semantic features of each basic block are extracted to generate a semantic vector. The basic blocks are clustered based on semantic vectors to obtain a set of semantic equivalence classes, and the main anchor points are identified based on the semantic equivalence classes and the stability of the basic blocks. Based on the main anchor point, construct the anchor point relationship diagram and determine the main edges to form the main anchor point skeleton diagram; The semantic core graph is obtained by semantically compressing and noise processing the control flow graph using the main anchor skeleton graph as a constraint. Perform hierarchical deterministic normalization on the semantic core graph and the main anchor skeleton graph, and generate a composite hash signature; The composite hash signature of the sample to be tested is matched with a pre-built malicious code family signature library to complete the determination of malicious code variants.

[0006] Optionally, constructing the control flow graph and extracting the semantic features of each basic block includes: Static disassembly is performed on the target executable file to identify function boundaries. Basic block nodes are divided based on instruction boundaries, and candidate control flow edge sets are obtained by analyzing instructions that involve indirect jumps or dynamically calculated target addresses, forming the control flow graph of the function. The instruction sequence within a basic block is traversed, and the distribution of instruction types is statistically analyzed to obtain opcode distribution characteristics. System interface calls and their parameter characteristics are identified to obtain API call characteristics. Constant and string information are extracted, and jump and branch structures are analyzed to obtain local control flow characteristics. Read and write role characteristics are extracted based on register and memory access patterns, and micro-semantic transfer characteristics are extracted based on the changes in the abstract state after instruction execution. The above-mentioned features are normalized and combined to generate the semantic vector of the corresponding basic block.

[0007] Optionally, the step of clustering basic blocks based on semantic vectors to obtain a set of semantic equivalence classes, and identifying the main anchor point based on the semantic equivalence classes and the stability of basic blocks, includes: Construct a family of locality-sensitive hash functions to perform hash mapping on the semantic vectors of each basic block, group basic blocks with the same hash code into the same cluster group, and perform semantic similarity verification on the basic blocks within the cluster group to obtain an initial set of semantic equivalence classes. The semantic consistency between each basic block and the center of its semantic equivalence class, the topological stability in the control flow graph, the behavioral semantic carrying capacity, and the noise probability are calculated. The basic block stability score is obtained by weighted calculation. The basic block with the highest stability score and the highest behavioral information content is marked as the main anchor point. When there are multiple main anchor points in the same semantic equivalence class, the basic block with the highest stability score is retained as the main anchor point.

[0008] Optionally, the step of constructing an anchor point relationship graph based on the main anchor point and determining the main edges to form a main anchor point skeleton graph includes: Construct an initial anchor point relationship graph using the identified main anchor points as nodes; Determine the path consistency, semantic continuity, reachability stability, and bridging noise ratio between any two main anchor points. Calculate the backbone score of candidate edges between anchor points, select candidate edges whose backbone scores meet the conditions as backbone edges, and form the main anchor point skeleton graph together with the selected backbone edges.

[0009] Optionally, the semantic compression and noise processing of the control flow graph based on the main anchor skeleton graph as constraints includes: Each semantic equivalence class is mapped to a compressed node, and jump edges between compressed nodes are established based on the original control flow edge relationships to form an initial compressed graph; During the compression process, anchor point constraint rules are executed. If the compressed node corresponding to the main anchor point does not meet the semantic equivalence class merging condition and does not destroy the main anchor point skeleton graph structure, it will not be merged with non-anchor point low stability nodes. Bridge nodes located between main anchor points and with stability below the preset threshold will be included in the noise buffer. Determine whether removing a node from the noise buffer affects the reachability between the main anchors. Mark nodes that do not affect the reachability as removable and absorbable. Determine whether a node only serves as a single path transfer. Mark nodes that only serve as transfers and have low semantic information as collapsible and absorbable. Retain nodes that carry independent behavioral semantics as ordinary compressed nodes to complete the noise buffering process.

[0010] Optionally, after noise buffering is completed, a topological consistency check is performed on the compressed graph to obtain the semantic core graph: Traverse all nodes and edges in the compressed graph and check for the existence of self-loop structures; Traverse the node pairs in the compressed graph and check for bidirectional conflicting edges; When a self-loop or contradictory edge is detected, the corresponding strongly connected component is located and clustering is performed. The above detection and merging steps are iteratively executed until no self-loops or contradictory edges appear in the compressed graph, thus obtaining the semantic core graph.

[0011] Optionally, performing hierarchical deterministic normalization on the semantic core graph and the main anchor skeleton graph and generating a composite hash signature includes: Calculate the corresponding semantic fingerprint for each compressed node in the semantic core graph, calculate the anchor fingerprint for the main anchor point in the main anchor point skeleton graph, and calculate the edge fingerprint for the backbone edge. According to the preset deterministic sorting rules, the anchor skeleton string and the ordinary compressed node supplement string are generated. The anchor skeleton string and the ordinary node supplement string are concatenated according to the preset weights and combined with the statistical information of the semantic core graph to form the final composite string. A composite hash signature is obtained by calculating the composite canonical string using a preset hash algorithm.

[0012] Optionally, matching the composite hash signature of the sample to be tested with the signature database includes: A two-level matching strategy is adopted. First, fast indexing and comparison are performed on the anchor skeleton specification string of the sample to be tested. If the skeleton matching passes, a fine-grained consistency check is then performed on the supplementary string and the semantic core graph statistics. Based on the two-level matching results, output whether it is a malicious code variant and its corresponding family tag.

[0013] Secondly, this application provides a malicious code variant detection system based on control flow graph semantic matching, comprising: The disassembly and feature extraction module is used to disassemble the target executable file, construct the control flow graph, extract the semantic features of each basic block, and generate semantic vectors. The semantic clustering and anchor point identification module is used to cluster basic blocks based on semantic vectors to obtain a set of semantic equivalence classes, and to identify the main anchor point based on the semantic equivalence classes and the stability of basic blocks; The skeleton graph construction module is used to construct an anchor point relationship graph based on the main anchor point and determine the main edges to form a main anchor point skeleton graph. The processing module is used to perform semantic compression and noise processing on the control flow graph with the main anchor skeleton graph as a constraint to obtain the semantic core graph; and to perform hierarchical deterministic normalization on the semantic core graph and the main anchor skeleton graph to generate a composite hash signature. The matching and determination module is used to match the composite hash signature of the sample to be tested with a pre-built malicious code family signature library to complete the determination of malicious code variants.

[0014] Thirdly, this application provides an electronic device, the electronic device comprising: At least one processor; and a memory communicatively connected to said at least one processor; The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method described in any one of the first aspects.

[0015] Compared with the closest prior art, the beneficial effects of this application are: This application proposes a method, system, and device for detecting malicious code variants based on control flow graph semantic matching. By introducing main anchor points and backbone edges, a stable skeleton spanning deformations is constructed to fundamentally resist interference factors such as bridging blocks, scheduling blocks, and pseudo anchor points. Simultaneously, anchor point constraint compression and noise delay absorption mechanisms are employed to effectively avoid semantic pollution and ensure the stable convergence of the semantic core graph.

[0016] This application employs a layered normalization and composite signature strategy, significantly improving signature invariance and ensuring high consistency among variants of the same origin. The two-level matching mechanism balances detection speed and robustness, with the anchor skeleton directly mapping malicious behavior chains, providing good interpretability and forensic capabilities.

[0017] This application breaks through the traditional approach of unified compression and unified standardization, and forms a three-stage convergence mode characterized by anchor point priority. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0019] Figure 1 This is a flowchart of a malicious code variant detection method based on control flow graph semantic matching provided in this application; Figure 2 This is a schematic diagram of the structure of a malicious code variant detection system based on control flow graph semantic matching provided in this application; Figure 3 This is a diagram of the internal structure of the electronic device provided in this application. Detailed Implementation

[0020] The embodiments of the technical solution of this application will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of this application and are therefore merely examples, and should not be construed as limiting the scope of protection of this application.

[0021] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains.

[0022] This application provides a method, system, and device for detecting malicious code variants using control flow graph semantic matching, applicable to the detection of malicious code variants using control flow graph semantic matching in complex obfuscated environments. The embodiments of this application are described below with reference to the accompanying drawings.

[0023] Example 1: As Figure 1 As shown, Embodiment 1 of this application provides a method for detecting malicious code variants by controlling flow graph semantic matching. This method specifically includes the following steps: S101 disassembles the target executable file, constructs a control flow graph, extracts the semantic features of each basic block, and generates a semantic vector. S102 clusters basic blocks based on semantic vectors to obtain a set of semantic equivalence classes, and identifies the main anchor point based on the semantic equivalence classes and the stability of basic blocks; S103 constructs an anchor point relationship diagram based on the main anchor point and determines the main edges to form a main anchor point skeleton diagram; S104 performs semantic compression and noise processing on the control flow graph using the main anchor skeleton graph as a constraint to obtain the semantic core graph; hierarchical deterministic normalization is performed on the semantic core graph and the main anchor skeleton graph to generate a composite hash signature; S105 matches the composite hash signature of the sample to be tested with a pre-built malicious code family signature library to complete the determination of malicious code variants.

[0024] The above step S101, which constructs the control flow graph and extracts the semantic features of each basic block, includes: Static disassembly is performed on the target executable file to identify function boundaries. Basic block nodes are divided based on instruction boundaries, and candidate control flow edge sets are obtained by analyzing instructions that involve indirect jumps or dynamically calculated target addresses, forming the control flow graph of the function. The instruction sequence within a basic block is traversed, and the distribution of instruction types is statistically analyzed to obtain opcode distribution characteristics. System interface calls and their parameter characteristics are identified to obtain API call characteristics. Constant and string information are extracted, and jump and branch structures are analyzed to obtain local control flow characteristics. Read and write role characteristics are extracted based on register and memory access patterns, and micro-semantic transfer characteristics are extracted based on the changes in the abstract state after instruction execution. The above-mentioned features are normalized and combined to generate the semantic vector of the corresponding basic block.

[0025] In one embodiment, in step S101 above, which involves constructing the control flow graph and extracting the semantic features of each basic block, a PE executable file under the Windows platform is used as an example. Static disassembly is performed using disassemblers such as IDA Pro or Ghidra. First, function boundaries are identified: by analyzing the executable file's entry point, function preambles (push ebp; mov ebp, esp, etc.), and calling conventions, the start and end addresses of each function are determined. For each function, basic blocks are divided according to instruction order: starting from the function's start instruction, the block is scanned sequentially. When a branch instruction (such as jmp, jz, call), return instruction (ret), or indirect jump instruction is encountered, the current basic block terminates, and a new basic block begins from the next instruction. For indirect jumps (such as jmp eax) or instructions that dynamically calculate target addresses, analysis is performed (e.g., "generating a superset of candidate edges"): the set of possible target addresses is recorded (e.g., inferred from the possible values ​​of global variables and registers), and all possible targets are added to the edge set as candidate control flow edges. The final result is a function-level control flow graph G=(V,E); where V is a basic block node and E is a control flow edge.

[0026] For each basic block v i Extract six types of semantic features: opcode distribution characteristics o i: Count the frequency of various instructions (mov, add, call, cmp, jcc, etc.) within a basic block and normalize them into a vector.

[0027] API call characteristics a i : Identify system APIs invoked by the call instruction (such as CreateFile, WriteProcessMemory), record the API name and parameter constants, and generate one-hot encodings or bag-of-words vectors.

[0028] Constants and string characteristics c i Extract numerical constants (e.g., 0x1234) and strings (e.g., “cmd.exe”) that appear in the instructions, and obtain features after hashing or word embedding.

[0029] Local control flow characteristics t i Analyze whether the basic block contains conditional jumps, loop structures, etc., and record the number of branches, jump offset range, etc.

[0030] Register / Memory Read / Write Role Characteristics r i For each instruction, mark the register and memory addresses it reads and writes, count the read and write patterns, and form a character feature vector.

[0031] Micro-semantic transfer features m i It employs abstract interpretation or symbolic execution to simulate the abstract state changes after instruction execution (such as stack pointer offset and flag changes), and encodes them as transfer vectors.

[0032] After performing Z-score normalization or minimum-maximum value normalization on the above six types of features respectively, they are concatenated to form a complete semantic vector s. i =[o i , a i , c i , t i , r i , m i ].

[0033] Step S102 above clusters the basic blocks based on semantic vectors to obtain a set of semantic equivalence classes, and identifies the main anchor points based on the semantic equivalence classes and the stability of the basic blocks, including: Construct a family of locality-sensitive hash functions to perform hash mapping on the semantic vectors of each basic block, group basic blocks with the same hash code into the same cluster group, and perform semantic similarity verification on the basic blocks within the cluster group to obtain an initial set of semantic equivalence classes. The semantic consistency between each basic block and its semantic equivalence class center, its topological stability in the control flow graph, its behavioral semantic carrying capacity, and its noise probability are calculated. A weighted average is then used to obtain a basic block stability score. Basic blocks with the highest stability score and the highest behavioral information content are marked as primary anchors. When multiple primary anchors exist within the same semantic equivalence class, the basic block with the highest stability score is retained as the primary anchor. The preset stability score condition can be set to a stability score greater than or equal to a preset stability threshold.

[0034] In one embodiment, step S102 above clusters basic blocks based on semantic vectors, using Locality Sensitive Hashing (LSH) for approximate nearest neighbor clustering: L hash functions h_1,...,h_L are selected, each function mapping the semantic vector to an integer. For each basic block... Calculate its hash vector H(s) i )=[h_1(s i ),...,h_L(s i All basic blocks with the same hash value vector are grouped into the same candidate cluster group. For basic blocks within each group, the cosine similarity or Euclidean distance between each pair is calculated. If the similarity exceeds a threshold (e.g., 0.85), they are merged into a semantic equivalence class. This yields the initial set of semantic equivalence classes {C_k}.

[0035] A semantic equivalence class is a set of basic blocks whose semantic vector similarity exceeds a preset threshold.

[0036] Furthermore, the stability score for each basic block is calculated. ; in, Let be the cosine similarity between the basic block and the center vector of its equivalence class, with a value in the range [0,1]. For topological robustness, it is defined as the reciprocal of the size of the largest connected component of the remaining graph after the node is removed from the control flow graph (or using a variant of PageRank), reflecting the importance of the node in the structure. The behavioral information is calculated based on the types of API calls within the basic block, constant complexity, etc., and normalized to [0,1]. The noise sensitivity is determined by the proportion of meaningless instructions (such as nop) and unreachable branches in the basic block; α, β, γ, and δ are weight parameters, with typical values ​​of α=0.4, β=0.3, γ=0.2, and δ=0.1.

[0037] Set the score threshold θ anchor =0.7, will Basic blocks with an Ω value ≥ 0.7 and the highest behavioral information content in their equivalence class are identified as candidates for primary anchor points. If there are multiple candidates in the same equivalence class, the one with the highest Ω value is retained as the primary anchor point, and the rest are downgraded to ordinary nodes.

[0038] Step S103 above, which involves constructing an anchor point relationship diagram based on the main anchor point and determining the main edges to form a main anchor point skeleton diagram, includes: Construct an initial anchor point relationship graph using the identified main anchor points as nodes; Determine the path consistency, semantic continuity, reachability stability, and bridging noise ratio between any two main anchor points. Calculate the backbone score of candidate edges between anchor points, select candidate edges whose backbone scores meet the conditions as backbone edges, and form the main anchor point skeleton graph together with the selected backbone edges.

[0039] In one embodiment, constructing the anchor point relationship graph can be performed through the following example: Let the set of main anchor points A = {a1,...,a...} q Construct the anchor point relationship graph G. a =(A, E a ), where E a Contains the original control flow path between all anchor pairs (there may be multiple paths). For each anchor pair (a p , a q First, obtain the information from a. p to a q The set of all directed paths P = {p1, p2, ..., p m Based on this set, the following metrics are extracted: PathCons: This measure measures the structural similarity of all paths in terms of node sequences. It is calculated as follows: for every two paths in the path set, the longest common subsequence (LCS) length is calculated. The average LCS length of all path pairs is then divided by the average path length (the average number of nodes in all paths) to obtain a normalized path consistency index. Semantic Continuity (SemCarry): This reflects the semantic importance of the behavior of non-anchor nodes on the path. It is calculated by: statistically analyzing the behavioral information content of all non-anchor nodes on the path, taking their arithmetic mean, and normalizing it to [0,1]. Reachability stability (ReachStab): This measures the probability that anchor pairs remain reachable after removing low-stability nodes. This patent uses a simplified calculation method: First, nodes on the path with stability scores below a preset threshold θ_{low}=0.4 are identified as low-stability nodes. Bridge Noise: Measures the relative proportion of low-stability nodes on the path.

[0040] Based on the above four indicators, the anchor pair (a) is calculated. p , a q The main score of the candidate edges between them is given by the following formula: Π(e_pq)=λ1·PathCons+λ2·SemCarry+λ3·ReachStab λ4·BridgeNoise; Wherein, λ1, λ2, λ3, and λ4 are weighting coefficients, all ranging from [0,1] and satisfying λ1+λ2+λ3+λ4=1. In this embodiment, typical values ​​are λ1=0.3, λ2=0.2, λ3=0.3, and λ4=0.2.

[0041] Set a threshold θ_trunk = 0.6, and determine edges with Π(e_pq) ≥ 0.6 as trunk edges. All trunk edges and the trunk edges together form the trunk skeleton graph G. a .

[0042] In step S104 above, the semantic compression and noise processing of the control flow graph using the main anchor skeleton graph as a constraint includes: Each semantic equivalence class is mapped to a compressed node, and jump edges between compressed nodes are established based on the original control flow edge relationships to form an initial compressed graph; During the compression process, anchor point constraint rules are executed. If the compressed node corresponding to the main anchor point does not meet the semantic equivalence class merging condition and does not destroy the main anchor point skeleton graph structure, it will not be merged with non-anchor point low-stability nodes. Bridge nodes located between main anchor points and with stability below a preset threshold are included in the noise buffer. The noise buffer is a logical area for temporarily storing low-stability nodes to be processed. Determine whether removing a node from the noise buffer affects the reachability between the main anchors. Mark nodes that do not affect the reachability as removable and absorbable. Determine whether a node only serves as a single path transfer. Mark nodes that only serve as transfers and have low semantic information as collapsible and absorbable. Retain nodes that carry independent behavioral semantics as ordinary compressed nodes to complete the noise buffering process.

[0043] In one embodiment, the above-mentioned semantic compression and noise processing under anchor constraints specifically involves: compressing each semantic equivalence class C... k Mapped to a compressed node u k Initial compressed image G c In the original control flow graph, if there exists a flow from C... i From any node in C j An edge to any node in u is then... i with u j Add directed edges between them.

[0044] Anchor point constraint rules are introduced: Compacted nodes corresponding to the main anchor point are prohibited from being directly merged with non-anchor point low-stability nodes, provided that the equivalence condition (i.e., the equivalence classes represented by the two compacted nodes can be merged) is not met and the order of the main anchor point's backbone edges is not disrupted. Specifically, if the main anchor point node u... p With non-anchor node u q There is an edge between them, but u q If the stability score is below 0.3, then u is not allowed. q Merge into u p The equivalence class of.

[0045] Bridge nodes located between two primary anchor points and with a stability score below 0.4 are placed in the candidate noise buffer. The following processing is performed on the nodes within the buffer: If deleting the node does not change the reachability between the two main anchors (there is an alternative path), it is marked as "deletable absorption"; if the node only appears on a single path and does not contain independent API calls or important constants, it is marked as "collapseable absorption" and merged into the adjacent anchor node; otherwise, it is retained as a normal compressed node.

[0046] This step, obtaining the semantic core graph, also includes performing a topological consistency check on the compressed graph: Traverse all nodes and edges in the compressed graph and check for the existence of self-loop structures; Traverse the node pairs in the compressed graph and check for bidirectional conflicting edges; When a self-loop or contradictory edge is detected, the corresponding strongly connected component is located and clustering is performed. The above detection and merging steps are iteratively executed until no self-loops or contradictory edges appear in the compressed graph, thus obtaining the semantic core graph.

[0047] Specifically, for the initial compressed graph G c Perform self-loop detection: Traverse each node and check if there is an edge pointing to itself. If a self-loop exists, merge the node with itself (i.e., there is a cycle within the equivalence class, and it needs to be checked whether it should be split into finer equivalence classes). Perform contradictory edge detection: If both u and u exist simultaneously... i →u j and u j →u i If these two nodes are contradictory, a bidirectional contradictory edge is formed, and the strongly connected component (SCC) containing these two nodes is located. All nodes within the SCC are merged into a new compressed node, and the edge relationships are updated. This process is repeated until there are no more self-loops or contradictory edges in the graph, resulting in the semantic core graph G. c .

[0048] Perform hierarchical deterministic normalization on the semantic core graph and the main anchor skeleton graph, and generate a composite hash signature; Furthermore, in step S104 above, performing hierarchical deterministic normalization on the semantic core graph and the main anchor skeleton graph and generating a composite hash signature includes: calculating the corresponding semantic fingerprint for each compressed node in the semantic core graph, calculating the anchor fingerprint for the main anchor in the main anchor skeleton graph, and calculating the edge fingerprint for the backbone edge. According to the preset deterministic sorting rules, the anchor skeleton string and the ordinary compressed node supplement string are generated. The anchor skeleton string and the ordinary node supplement string are concatenated according to the preset weights and combined with the statistical information of the semantic core graph to form the final composite string. A composite hash signature is obtained by calculating the composite canonical string using a preset hash algorithm.

[0049] The preset deterministic sorting rules include: main anchor points are sorted according to their topological order in the main anchor point skeleton graph; nodes at the same level are arranged in lexicographical order of their node identifiers; and ordinary compressed nodes are arranged in ascending order of their node identifiers. The anchor point skeleton specification string is generated by sequentially concatenating "anchor point fingerprint + outgoing edge fingerprint".

[0050] Specifically, for each compressed node u in the semantic core graph k Calculate the semantic fingerprint F(u) k =SHA-256(sort({s i |v i ∈C k This involves sorting all the semantic vectors of the basic blocks corresponding to the node, concatenating them, and then hashing them. Here, `sort` indicates a deterministic sorting of the set elements in lexicographical order. For the primary anchor node, calculate the anchor fingerprint A(u) k )=SHA-256(F(u k )|Role(u k )|Rank_Ga(u k ), where SHA-256 represents a 256-bit secure hash algorithm; Role(u k ) marks whether the anchor point is of inlet / outlet type, Rank_Ga(u k () represents the sorting number in the skeleton diagram. For the main edge e... pq The edge fingerprint is calculated using the following formula: B(e pq )=SHA-256(A(u p )|A(u q )|Type(e pq Type(e) pq ) represents the main edge e pq The type (sequence, branching, looping, etc.). A(u p ), A(u) q ) represent the main anchor node u respectivelyp With non-anchor node u q Anchor fingerprint.

[0051] Arrange all primary anchor points in topological order, and output the anchor point fingerprints and outgoing edge fingerprints sequentially to form the anchor point skeleton canonical string S. a For all ordinary compressed nodes that are not primary anchors, sort them by node identifier and output their semantic fingerprints to form the supplementary string S. c The statistical semantic core graph is analyzed to obtain the statistical summary Stat(G). Information such as the number of nodes, the number of edges, and the average out-degree are collected. c Final compound signature (Sign) + =SHA-256(S a |η·S c |ζ·Stat(G c )), η and ζ are weighting coefficients (e.g., η=2, ζ=1, indicating S c The number of times the string "Stat" is repeated.

[0052] Step S105 above, matching the composite hash signature of the sample to be tested with the signature database, includes: A two-level matching strategy is adopted. First, fast indexing and comparison are performed on the anchor skeleton specification string of the sample to be tested. If the skeleton matching passes, a fine-grained consistency check is then performed on the supplementary string and the semantic core graph statistics. Based on the two-level matching results, output whether it is a malicious code variant and its corresponding family tag.

[0053] In one embodiment, a malware family signature library F is pre-built. f = {Sign1, Sign2, …}, where each family corresponds to one or more composite signatures (which can be generated from multiple representative samples). During detection, the anchor skeleton canonical string S of the sample to be tested is first extracted. a A fast comparison is performed in the signature database index (e.g., using a hash table). If a matching family candidate is found, the complete composite signature (Sign) is then calculated. + The signature is precisely compared with that of the candidate family. If they are completely identical or the Hamming distance is below the threshold, the sample is determined to be a variant of the malicious code in that family; otherwise, it is marked as unknown or benign.

[0054] The pre-built malware family signature library is obtained by performing steps S101 to S104 on representative samples of known families to generate composite hash signatures, which are then stored according to family labels; the fast index comparison is implemented using a hash table or inverted index structure.

[0055] Example 2: The following example of detecting a variant of a certain Trojan family will be used to further illustrate this solution.

[0056] Obtain the original sample of this family and two variant samples. Variant D uses control flow flattening technology to reorganize multiple functional modules that were originally executed sequentially into a distribution structure centered on the scheduler. Variant E performs surface rewriting of key behavior blocks through equivalent instruction replacement and register renaming.

[0057] First, static disassembly was performed on the three samples to construct function-level control flow graphs. Complete semantic vectors were extracted for each basic block, including opcode distribution features, application programming interface (API) call features, constant and string features, local control flow features, register / memory read / write role features, and micro-semantic transfer features. For key blocks involving network connection establishment, remote command reception, local information collection, and data transmission, their micro-semantic transfer features exhibit a composite pattern of external interaction, data transport, and conditional decision-making, while the register / memory read / write role features show a normalized pattern of parameter reading and buffer writing.

[0058] After initial clustering using locality-sensitive hashing, a stability score is calculated for each basic block. Several basic blocks involved in network initialization, command parsing, information gathering, and data transmission, with high semantic consistency, strong topological stability, and high behavioral semantic carrying capacity, are selected as primary anchors. However, a large number of distribution blocks introduced by the flattened scheduler in variant D have high noise probability and stability scores below the stability threshold, and are not selected as anchors.

[0059] A main anchor skeleton graph is constructed based on the reachability relationships between main anchors. Stable backbone behavior chains are identified: the network initialization anchor points to the command parsing anchor, the command parsing anchor points to the information gathering anchor, and the information gathering anchor points to the data transmission anchor. A backbone score is calculated for each candidate anchor edge, and edges with consistent direction and stable semantic continuity are retained as backbone edges. Multiple equivalent distribution paths generated by the flattened scheduler in variant D are not included in the skeleton graph due to their high proportion of bridging noise and insufficient backbone score.

[0060] During graph compression, anchor point constraints are applied: primary anchor points cannot be directly merged with ordinary low-stability blocks; low-stability bridging blocks located between primary anchor points enter the candidate noise buffer. For distribution blocks introduced by flattening in variant D, removal does not affect the reachability and order relationships between primary anchor point pairs and is marked as removable; for a few distribution blocks that undertake independent transit semantics, they are folded and absorbed into the edge attributes of adjacent backbone edges. After compression, the three samples converge to the same primary anchor point skeleton.

[0061] A composite signature is generated during the normalization phase. The anchor skeleton sub-signatures of the original sample are completely identical to those of variants D and E. During detection, the corresponding Trojan family is quickly identified through primary skeleton matching, and the variant affiliation is confirmed through secondary supplementary matching. If the primary matching is not complete but the anchor skeleton edit distance is below the suspected threshold, manual review is triggered.

[0062] Example 3: The following example, using the detection of a certain worm family variant, further illustrates this scheme.

[0063] Obtain the original sample of the family and two variant samples. Variant F splits the original complete infection logic into multiple fine-grained blocks through basic block splitting, while variant G adds a large number of irrelevant jumps on key behavior paths through dead code insertion.

[0064] Static disassembly was performed on the three samples to construct function-level control flow graphs. Semantic vectors were extracted from each basic block, with particular attention paid to key blocks involving network scanning, vulnerability exploitation, payload delivery, and self-replication. The micro-semantic transfer features of these key blocks exhibited a composite pattern of environment detection, external interaction, and control distribution.

[0065] After initial clustering using locality-sensitive hashing, stability scores are calculated. Several basic blocks involved in network scanning, vulnerability exploitation, load balancing, and self-replication are selected as primary anchors; while multiple low-information intermediate blocks splitting from variant F and dead code blocks inserted in variant G have high noise probabilities and stability scores below the stability threshold, and are not selected as anchors.

[0066] Construct the main anchor point skeleton graph and identify the backbone behavior chain: network scanning anchor points to exploit anchor points, exploit anchor points to load balancing anchor points, and load balancing anchor points to self-replication anchor points. Calculate the backbone score for candidate anchor point edges and retain stable edges as backbone edges.

[0067] Anchor point constraint rules are applied during graph compression. For low-information intermediate blocks split from variant F, if removal does not affect the reachability and order relationships between main anchor pairs, they are marked as removable and absorbable; if they only serve as a single path transition and have low semantic information, they are folded and absorbed into the edge attributes of adjacent main edges. For dead code blocks inserted in variant G, if removal does not affect the reachability relationships between main anchor pairs, they are marked as removable and absorbable. After compression, the three samples converge to the same main anchor skeleton.

[0068] A composite signature is generated and a two-level matching test is performed. The anchor skeleton sub-signatures of the original sample are completely consistent with those of variants F and G. The worm family is identified through the first-level skeleton matching, and the variant affiliation is confirmed through the second-level supplementary matching.

[0069] Example 4: Based on the same technical concept, Example 4 of this application also provides a malicious code variant detection system based on control flow graph semantic matching, such as... Figure 2 As shown, it includes: a disassembly and feature extraction module 210, a semantic clustering and anchor point recognition module 220, a skeleton graph construction module 230, a processing module 240, and a matching determination module 250. Intermediate results are exchanged between modules through standardized data interfaces (such as JSON or Protocol Buffers), and it can be deployed on a server or cloud security analysis platform. Among them: The disassembly and feature extraction module 210 is used to disassemble the target executable file, construct a control flow graph, extract the semantic features of each basic block, and generate a semantic vector. The semantic clustering and anchor point identification module 220 is used to cluster basic blocks based on semantic vectors to obtain a set of semantic equivalence classes, and to identify the main anchor point based on the semantic equivalence classes and the stability of basic blocks; The skeleton graph construction module 230 is used to construct an anchor point relationship graph based on the main anchor point and determine the main edges to form a main anchor point skeleton graph. Processing module 240 is used to perform semantic compression and noise processing on the control flow graph with the main anchor skeleton graph as a constraint to obtain a semantic core graph; and to perform hierarchical deterministic normalization on the semantic core graph and the main anchor skeleton graph to generate a composite hash signature. The matching and determination module 250 is used to match the composite hash signature of the sample to be tested with the pre-built malicious code family signature library to complete the determination of malicious code variants.

[0070] Example 5: In one embodiment, Example 5 of this application also provides an electronic device; the electronic device may be a terminal, and its internal structure diagram may be as follows. Figure 3As shown. The electronic device includes a processor, memory, communication interface, display screen, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements the malicious code variant detection method based on control flow graph semantic matching as described in any one of steps S101 to S105. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input device can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.

[0071] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0072] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0073] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0074] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0075] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0076] The above are merely embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application shall be included within the scope of the claims of this application pending approval.

Claims

1. A method for detecting malicious code variants based on control flow graph semantic matching, characterized in that, include: The target executable file is disassembled, a control flow graph is constructed, and the semantic features of each basic block are extracted to generate a semantic vector. Basic blocks are clustered based on semantic vectors to obtain a set of semantic equivalence classes. The main anchor points are then identified based on the semantic equivalence classes and the stability of the basic blocks, including: Construct a family of locality-sensitive hash functions to perform hash mapping on the semantic vectors of each basic block, group basic blocks with the same hash code into the same cluster group, and perform semantic similarity verification on the basic blocks within the cluster group to obtain an initial set of semantic equivalence classes. Calculate the semantic consistency between each basic block and the center of its semantic equivalence class, the topological robustness in the control flow graph, the behavioral semantic carrying capacity, and the noise sensitivity. Perform weighted calculation to obtain the basic block stability score. Mark the basic block with the highest stability score and the highest behavioral semantic carrying capacity as the main anchor point. When there are multiple main anchor points in the same semantic equivalence class, retain the basic block with the highest stability score as the main anchor point. The noise sensitivity is determined by the proportion of meaningless instructions and unreachable branches in the basic block; Based on the main anchor point, an anchor point relationship graph is constructed and the main edges are determined to form a main anchor point skeleton graph, including: Construct an initial anchor point relationship graph using the identified main anchor points as nodes; Obtain all directed paths between any two main anchor points, and calculate path consistency, semantic continuity, reachability stability, and bridging noise ratio based on all directed paths; The semantic continuity is used to reflect the semantic importance of the behavior of non-anchor nodes on the path, the reachability stability is used to measure the probability that anchor pairs remain reachable after deleting low-stability nodes, and the bridging noise ratio is used to measure the relative proportion of low-stability nodes on the path. The path consistency, semantic continuity, reachability stability, and the proportion of bridging noise in the path are weighted and summed to obtain the backbone score of the candidate edges between anchor points. Candidate edges whose backbone scores meet the conditions are selected as backbone edges. The main anchor points and the selected backbone edges together form the main anchor point skeleton graph. The control flow graph is semantically compressed and noise-processed using the main anchor skeleton graph as a constraint to obtain a semantic core graph, including: mapping each semantic equivalence class to a compressed node, establishing jump edges between compressed nodes according to the original control flow edge relationships, and forming an initial compressed graph; Anchor point constraint rules are executed during the formation of the initial compressed graph; Perform hierarchical deterministic normalization on the semantic core graph and the main anchor skeleton graph to generate a composite hash signature, including: Calculate the corresponding semantic fingerprint for each compressed node in the semantic core graph, calculate the anchor fingerprint for the main anchor point in the main anchor point skeleton graph, and calculate the edge fingerprint for the backbone edge. Arrange all main anchor points according to their topological order in the main anchor point skeleton graph, and output the anchor point fingerprint of each main anchor point and the edge fingerprint of its outgoing edge in sequence to form the anchor point skeleton specification string. After sorting all non-main anchor point ordinary compressed nodes by node identifier, output their semantic fingerprints to form supplementary strings for ordinary compressed nodes; The anchor skeleton specification string and the ordinary compressed node supplementary string are concatenated according to preset weights, and combined with the statistical information of the semantic core graph to form the final composite specification string; A composite hash signature is obtained by calculating the composite canonical string using a preset hash algorithm. The composite hash signature of the sample to be tested is matched with a pre-built malicious code family signature library to complete the determination of malicious code variants.

2. The method according to claim 1, characterized in that, The construction of the control flow graph and extraction of semantic features of each basic block includes: Static disassembly is performed on the target executable file to identify function boundaries. Based on the function boundaries, basic block nodes are divided, and the set of candidate control flow edges is obtained by analyzing instructions that involve indirect jumps or dynamic calculation of target addresses, thus forming the control flow graph of the function. The instruction sequence within a basic block is traversed, and the distribution of instruction types is statistically analyzed to obtain opcode distribution characteristics. System interface calls and their parameter characteristics are identified to obtain API call characteristics. Constant and string information in the instructions are extracted to obtain constant and string characteristics. Jump and branch instructions are analyzed to obtain local control flow characteristics. Read and write role characteristics are extracted based on register and memory access patterns. Micro-semantic transfer characteristics are extracted based on the changes in the abstract state after instruction execution. The above-mentioned multiple features are normalized and combined to generate the semantic vector of the corresponding basic block.

3. The method according to claim 1, characterized in that, The anchor point constraint rule executed during the formation of the initial compressed graph is as follows: the compressed node corresponding to the main anchor point will not be merged with non-anchor point low-stability nodes if the semantic equivalence class merging condition is not met and the main anchor point skeleton graph structure is not destroyed; bridging nodes located between main anchor points and whose stability is lower than a preset threshold will be included in the noise buffer. Determine whether removing a node from the noise buffer affects the reachability between the main anchors, and mark nodes that do not affect the reachability as removable and absorbable. Determine whether a node only serves as a single-path transfer node. Nodes that only serve as single-path transfer nodes and have low semantic information content are marked as collapsible and absorbable. Nodes that carry independent behavioral semantics are retained as ordinary compressed nodes, thus completing the noise buffering process.

4. The method according to claim 3, characterized in that, After noise buffering is completed, the process also includes performing a topological consistency check on the initial compressed graph to obtain the semantic core graph: Traverse all nodes and edges in the initial compressed graph and check for self-loop structures; if a self-loop is detected, check whether it should be split into finer equivalence classes. Traverse the node pairs in the initial compressed graph and check for contradictory edges; contradictory edges refer to bidirectional directed edges; when a contradictory edge is detected, locate the corresponding strongly connected component and perform cluster merging; The above detection and merging steps are executed iteratively until no self-loops or contradictory edges appear in the initial compressed graph, thus obtaining the semantic core graph.

5. The method according to claim 1, characterized in that, The step of matching the composite hash signature of the sample to be tested with the signature database includes: A two-level matching strategy is adopted: First, the anchor skeleton canonical string of the sample to be tested is extracted and quickly compared in the index of the signature library; If a matching family candidate is found, the complete composite signature is further calculated and precisely compared with the signature of the candidate family. Based on the two-level matching results, output whether it is a malicious code variant and its corresponding family tag.

6. A malicious code variant detection system based on control flow graph semantic matching, employing the method described in claim 1, characterized in that, include: The disassembly and feature extraction module is used to disassemble the target executable file, construct the control flow graph, extract the semantic features of each basic block, and generate semantic vectors. The semantic clustering and anchor point identification module is used to cluster basic blocks based on semantic vectors to obtain a set of semantic equivalence classes, and to identify the main anchor point based on the semantic equivalence classes and the stability of basic blocks; The skeleton graph construction module is used to construct an anchor point relationship graph based on the main anchor point and determine the main edges to form a main anchor point skeleton graph. The processing module is used to perform semantic compression and noise processing on the control flow graph with the main anchor skeleton graph as a constraint to obtain the semantic core graph; and to perform hierarchical deterministic normalization on the semantic core graph and the main anchor skeleton graph to generate a composite hash signature. The matching and determination module is used to match the composite hash signature of the sample to be tested with a pre-built malicious code family signature library to complete the determination of malicious code variants.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to said at least one processor; The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Code homology analysis method and device

    CN118012480A

  • Network security malicious code binary search method and system

    CN121509118A