An automated vulnerability mining method for potential configuration attributes of internet of things devices

CN122457374BActive Publication Date: 2026-09-15NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610858565.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-09-15
Estimated Expiration
2046-06-15

AI Technical Summary

Technical Problem

与POSIX标准文件系统API不同,物联网厂商普遍采用高度定制的私有实现:部分配置读取函数依赖整数索引而非字符串键访问数据,另一部分通过多参数接口访问结构化或分层的配置数据;同时核心处理逻辑往往被封装于多层包装函数之内,且生产固件普遍经过符号剥离,使得基于字符串交叉引用、固定签名匹配或函数命名启发的方法难以稳定识别配置处理函数

Benefits of technology

[0017] This invention achieves systematic mining and utilization of potential configuration attributes—a long-neglected attack surface—during the vulnerability discovery process of IoT devices. Compared with existing technologies, this invention has the following significant advantages:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122457374B_ABST
    Figure CN122457374B_ABST
Patent Text Reader

Abstract

The application discloses an automatic vulnerability mining method for potential configuration attributes of Internet of Things equipment, constructs a function behavior pedigree graph across encapsulation layers by utilizing the behavior symmetry of configuration access operations, and locates configuration processing functions in a de-symbolized firmware binary based on a graph neural network; subsequently, through joint pruning of hierarchical strict post-dominance analysis and immutable data source constraints, volatile attributes determined to be reset in the startup phase are removed, and the false positive rate is significantly reduced; by introducing a configuration-oriented lightweight taint engine, the availability of configuration attributes is efficiently determined under the condition of a short Def-Use distance; finally, the image equipment is used as a semantic converter to realize transparent tampering and injection of encrypted configuration files. The application can systematically mine potential configuration attribute vulnerabilities that are invisible in the front end and actually processed in the back end in cross-vendor and cross-architecture IoT firmware, and fills the blind spot of existing vulnerability mining methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of IoT security and software vulnerability mining technology, specifically to an automated vulnerability mining method for potential configuration attributes of IoT devices. Background Technology

[0002] IoT devices have become deeply integrated into everyday scenarios such as smart homes, security monitoring, and industrial control, bringing great convenience to people's lives. According to publicly available statistics, the number of interconnected IoT devices worldwide is expected to exceed 22 billion by 2027. However, these devices are generally deployed in resource-constrained environments, running highly customized lightweight firmware, and lacking comprehensive security protection mechanisms. Recent security reports show that global IoT attacks continue to rise, with the annual attack scale showing a significant increase of over 40% compared to the previous year. Mirai and its variants have even launched unprecedented distributed denial-of-service attacks by exploiting firmware vulnerabilities in a massive number of devices. Gateway devices such as routers and network-attached storage (NAS) devices that expose web management services are particularly vulnerable—their backend logic directly receives user input, and any flaws can be remotely exploited, further becoming a springboard for infiltrating other devices on the internal network. Therefore, early and systematic vulnerability discovery in IoT devices is crucial for ensuring the security of terminals and networks.

[0003] In attack surface research on IoT web services, much work has focused on explicit parameters—vulnerabilities where security researchers can influence backend processing by controlling frontend form field inputs. However, this research paradigm has a significant blind spot: it ignores a set of configuration attributes that are actually processed by the firmware backend logic but are invisible to the frontend. These configuration attributes can be indirectly injected by modifying and uploading configuration files; this invention defines them as "latent configuration attributes." Firmware hardening trends make these hidden fields increasingly key entry points for attackers to cross privilege boundaries and achieve persistent control: a single configuration injection can complete vertical privilege escalation from a low-privilege user to root privileges, and these privileges are difficult to remove even after device reboots or firmware upgrades.

[0004] Existing vulnerability discovery efforts targeting IoT devices can be broadly categorized into three types, but none effectively cover the attack surface of potential configuration attributes: The first type is static taint analysis methods driven by the web front-end, focusing on tracing the data flow from front-end HTTP parameters to the back-end sink, inherently blinding to configuration fields invisible to the front-end; the second type is binary fuzzing methods, using network protocols or web interfaces as mutation entry points, unable to effectively mutate structured, encrypted configuration files, resulting in extremely limited coverage; the third type is semantic analysis methods based on large language models, which, while possessing strong semantic reasoning capabilities, suffer from high false positive rates and difficulty in reliably locating configuration processing functions in de-symbolized binary data when faced with highly customized configuration access semantics such as NVRAM and shared memory. In summary, none of these three methods can systematically identify potential configuration attributes, determine their exploitability, or effectively inject them into encrypted configuration files.

[0005] Furthermore, automating the analysis and utilization of potential configuration attributes in IoT devices faces three fundamental challenges:

[0006] (1) Obstacles to identifying configuration processing functions. Unlike the POSIX standard file system API, IoT manufacturers generally use highly customized proprietary implementations: some configuration reading functions rely on integer indices rather than string keys to access data, while others access structured or hierarchical configuration data through multi-parameter interfaces; at the same time, the core processing logic is often encapsulated in multiple layers of wrapper functions, and production firmware generally undergoes symbol stripping, making it difficult to reliably identify configuration processing functions using methods based on string cross-referencing, fixed signature matching, or function naming heuristics.

[0007] (2) The ambiguity of the availability verification of potential attributes. On the one hand, although some configuration fields are writable, they will be deterministically reset to hard-coded constants during the system initialization phase. No matter how the user tampers with them, they will be overwritten. These are called volatile attributes. On the other hand, automatically determining whether a non-volatile attribute can eventually flow to a sensitive sink point is also a core challenge in the verification phase.

[0008] (3) Injection barriers to encrypted configuration files. To ensure configuration security, manufacturers usually implement hash verification, compression and obfuscation, and high-strength encryption on configuration files. Any direct modification to the ciphertext will destroy the file structure or cause the integrity verification to fail, making it extremely difficult to inject legitimate configuration files without the key and algorithm details. Summary of the Invention

[0009] The purpose of this invention is to address the shortcomings of existing technologies and provide an automated vulnerability discovery method for potential configuration attributes of IoT devices. This method enables the systematic identification, exploitability verification, and cryptographic injection of potential configuration attributes that are not visible to the front end but are actually processed in the back end, thereby improving the ability of security personnel to discover deep vulnerabilities in IoT devices.

[0010] The technical solution to achieve the purpose of this invention is: an automated vulnerability discovery method for potential configuration attributes of Internet of Things (IoT) devices, comprising the following steps:

[0011] Step 1: Pruning candidate domains: Unpack the target firmware, identify the boundary binary related to system initialization and network services, and extract the dynamic link libraries that the binary files depend on; exclude standard library components through a whitelist mechanism, and narrow the analysis scope to the vendor's private shared libraries that carry core business logic, as candidate domains for locating configuration processing functions;

[0012] Step 2: Function Behavior Genealogy Construction: Perform lightweight data flow analysis on each function in the candidate private library and assign atomic behavior labels according to the data flow of parameters and return values; apply complexity gradient constraints and data flow convergence constraints on the call relationship, and aggregate semantically related encapsulated functions and core parent functions into a function behavior genealogy.

[0013] Step 3: Configuration processing function localization based on graph isomorphic network: Vectorize and embed each node in the function behavior spectrum graph using a graph isomorphic network with random parameters, and aggregate the full graph vectors using the inverse of the topological distance from the node to the parent function as the affinity weight; In the embedding space, the OPTICS clustering algorithm is used to identify structurally homogeneous nodes and filter outliers as candidate clusters; Based on the centrality index, the candidate clusters are ranked by their in-degree and reference frequency in the global call graph, and the dual families with the highest ranking are determined as the configuration write function family and the configuration read function family;

[0014] Step 4: Volatile Attribute Pruning: Promote the binary code to a normalized intermediate representation and build a system dependency graph. Treat all API call points where configurations are written as pollution points. Perform strict post-domination analysis within the process and reachability analysis between processes on each pollution point. Then, add immutable data source constraints and remove volatile attributes that are deterministically reset on the startup path.

[0015] Step 5: Lightweight taint analysis for configuration attributes: Using the configuration read function family as taint sources and dangerous library functions and command execution functions as taint functions, copy length constraints are recorded during propagation and the target buffer space is compared to determine buffer overflow reachability; the completeness of the purifier is checked on the command injection path based on a predefined purifier rule base; finally, the system key field knowledge base is used to filter out unusable attributes in the project, and the final list of potentially usable configuration attributes is output.

[0016] Step 6: Mirror Device Configuration Injection: Obtain a validator local device of the same model as the victim device, import the encrypted configuration file exported by the victim device and decrypt it; modify the exploitable potential configuration attributes in the non-volatile storage of the validator local device into the attack payload; call the built-in export interface of the validator local device to re-encapsulate and encrypt it to obtain a malicious configuration file with the same encryption format as the victim device, and then upload it to the victim device to complete the injection.

[0017] This invention achieves systematic mining and utilization of potential configuration attributes—a long-neglected attack surface—during the vulnerability discovery process of IoT devices. Compared with existing technologies, this invention has the following significant advantages:

[0018] (1) Complete attack surface coverage. This invention is the first to systematically reveal and automatically mine the attack surface composed of potential configuration attributes that are invisible to the front end but actually processed in the back end. It fills the common blind spot of the three existing methods: static taint analysis of Web front-end drivers, protocol-level fuzz testing and large model semantic analysis. It can discover deep vulnerabilities that traditional methods cannot reach.

[0019] (2) Strong generalization capability across vendors and architectures. This invention locates configuration processing functions based on the behavioral symmetry of configuration access operations rather than vendor-specific signatures. It has stable recognition capabilities for firmware that is de-symbolized, cross-instruction set architecture, and cross-storage backend (NVRAM key-value, integer index MIB table, XML configuration and proprietary binary format), without relying on any function naming heuristics or string cross-references.

[0020] (3) Low false positive rate. This invention uses a three-layer cascaded pruning system of "strict post-domination within the process L1 - reachability between processes L2 - immutable data source constraints" to accurately remove volatile attributes. Combined with a nine-layer taint filter chain and a knowledge base of key system fields to filter unusable paths in the process, the final output vulnerability list has a high true positive rate.

[0021] (4) Injection capability that breaks through encryption protection. The mirror device configuration injection method proposed in this invention takes advantage of the common implementation defect of sharing static encryption keys among devices of the same model. It uses the verifier's local device of the same model as a semantic converter to complete the legitimate encapsulation of the encrypted configuration file without reversing the encryption algorithm or extracting the key. It can achieve nearly 100% covert and persistent injection under the protection of encrypted configuration. Attached Figure Description

[0022] Figure 1 This is a system structure diagram of the present invention.

[0023] Figure 2 This is a basic flowchart of the present invention for automated vulnerability mining of potential configuration attributes of IoT devices.

[0024] Figure 3 This is a schematic diagram of the construction of the function behavior genealogy diagram FBLG of the present invention. Detailed Implementation

[0025] The present invention will now be described in further detail with reference to the accompanying drawings and examples. The following embodiments are implemented based on the technical solution of the present invention, providing detailed implementation methods and processes; however, the scope of protection of the present invention is not limited to the following embodiments.

[0026] An automated vulnerability discovery method targeting the potential configuration attributes of IoT devices includes the following steps:

[0027] Step 1: Candidate Domain Pruning

[0028] The target firmware image was subjected to automated unpacking tools to extract the root file system, and the system initialization scripts and startup instructions of daemons such as Web services, Telnet / SSH, and SNMP were analyzed to locate the boundary binary sets that carry user input. Extract all linked dynamic libraries by parsing ELF dependencies and library loading paths; based on the standard library whitelist, remove common basic libraries such as libc, libm, and libpthread, retaining only the set of private shared libraries that carry vendor-customized configuration access logic. This serves as the input for constructing the subsequent function behavior genealogy graph; the original boundary binary... It will then be retained as an entry candidate for the subsequent vulnerability verification stage.

[0029] Step 2: Constructing the genealogical diagram of function behavior

[0030] To alleviate the semantic dilution problem caused by multi-level call encapsulation, this invention proposes a Function Behavior Lineage Graph (FBLG) to characterize the propagation of the behavioral features of core configuration functions along the call chain to encapsulated functions. The specific process is as follows:

[0031] (1) Atomic behavior tagging: Perform lightweight data flow analysis on each function f and assign an atomic tag Tag(f). If the parameter data flow of f is eventually dumped into persistent storage or memory write operation, then Tag(f) = Set; if the return value or output parameter of f comes from persistent storage or global state, then Tag(f) = Get; for other functions, Tag(f) = Null.

[0032] (2) Inheritance constraint extension: Based on the premise of consistent atomic behavior marking, two types of constraints are introduced to ensure semantic consistency between the encapsulated function and the core function:

[0033] Complexity gradient constraint: requires that the cyclomatic complexity of the sub-function is strictly less than that of the parent function, i.e.

[0034]

[0035] The Complexity function is a cyclomatic complexity function. This constraint ensures that inheritance only extends to peripheral lightweight auxiliary interfaces or encapsulated logic, thereby eliminating intermediary functions that call core functions but contain a large amount of irrelevant business logic.

[0036] Data flow convergence constraint: This requires a clear bidirectional data dependency path between the wrapper function and its parent function—the return value of the wrapper function comes from the return value of the parent function, and the parameters of the parent function come from the wrapper function. This constraint ensures that tight semantic connectivity is maintained within the genealogy.

[0037] (3) Tail call optimization: For semantic copying scenarios caused by compiler tail call optimization, the encapsulated function is regarded as a semantic alias of the parent function, and is an independent node but inherits all feature vectors of the parent function.

[0038] After the above process, functions with the same atomic behavior markers in the candidate private library are organized into a family of functions rooted in the core parent function, forming the basic node unit of FBLG.

[0039] Step 3: Locating the configuration processing function based on graph isomorphic networks

[0040] (1) Function feature extraction: This invention constructs a multi-dimensional function representation system that integrates structural features, data flow features, and API behavior semantics. Structural features are extracted based on the topological information of the function control flow graph (CFG), data flow features are extracted based on the parameter dependency relationship of the data dependency graph (DDG), and the semantic level is characterized by the call distribution of key system APIs (such as network communication, file I / O, memory access, and string processing) to depict the function behavior pattern.

[0041] (2) Structure-aware embedding: Random-GIN, a three-layer graph isomorphic network with randomized parameters, is used to embed each node in FBLG into a vectorized form, resulting in a d=128-dimensional feature vector. To strengthen the dominant role of the core parent function in the overall features, this invention introduces a proximity-weighted mechanism: node weights. Topological distance from its parent function Inversely proportional, calculated using the following formula:

[0042]

[0043] Based on this, the full-image feature vector of FBLG is weighted and aggregated:

[0044]

[0045] (3) Dual Family Matching and Ranking: The OPTICS density clustering algorithm is used to cluster function vectors in the embedding space to identify structurally homologous function clusters and filter outlier nodes as candidate clusters; semantic completeness constraints are imposed on each candidate cluster—it must contain both Set-Tag and Get-Tag nodes to be considered a potential CIF dual family; further, based on the use of centrality index (The sum of the in-degree and reference frequency of potential CIF candidate dual families in the global call graph), sort all candidate dual families according to the following formula, and take the Top-1 as the target dual family:

[0046]

[0047] Step 4: Volatility Attribute Pruning

[0048] To eliminate volatile properties that are deterministically reset during system initialization, this invention first elevates the binary code to a normalized intermediate representation and constructs a system dependency graph SDG = (V, E), where V represents basic blocks and function call points, and E integrates control flow, data flow, and call graph information; the set of all configuration calls written to the API is defined as the sink point set. For each The following strict post-dominance analysis was performed stratified:

[0049] L1 process post-domination: For a function f containing a sink point s, construct the control flow graph. Given a post-dominant tree, determine if s is an entry node. Strictly dominant nodes, i.e.

[0050]

[0051] If this condition is true, it proves that as long as function f is executed, operation s will necessarily be triggered.

[0052] L2 inter-procedural reachability: Further backtracking at the call graph level from the initialization entry point Root. boot The call chain π to function f = ⟨Root boot , …, caller i The function f> verifies whether each call point in the chain satisfies the L1 post-domination condition of its parent function. When all call points in the chain satisfy the strict post-domination relationship, the transitivity of the domination relationship indicates that operation s has global determinism during system startup.

[0053] Immutable data source constraint: Even if the control flow is deterministic, if the write parameter val comes from a previous read (e.g., "v= nvram_get(k); nvram_set(k, v);"), it is a state recovery rather than a true reset. Therefore, this invention adds a data flow constraint: when performing a backward slice on the parameter val at sink point s, the slice must terminate at a static constant in the .rodata segment to be considered a true volatile write.

[0054] By combining control flow and data flow constraints, this step can effectively eliminate volatile or recoverable configuration writes, significantly reducing false alarms in subsequent stages.

[0055] Step 5: Lightweight taint analysis for configuration attributes

[0056] This step is based on step 3. Using dangerous library functions and command execution functions as taint sources, and dangerous library functions and command execution functions as taint sinks, the reachability of configuration attributes to sink points is traced in the firmware binary. Based on the "just-in-time" behavior pattern of configuration attributes (i.e., backend programs typically read configuration values ​​only close to critical logic, resulting in a significantly shorter definition-use distance), this invention designs the following dedicated purifier:

[0057] (1) Buffer overflow detection: The length constraint is dynamically recorded each time a data transfer occurs. For example, when passing through strncpy(v5, v11, 80), the maximum write length of v5 is recorded as constrained to 80. When the taint finally reaches the sink, the most stringent constraint will be accumulated. With target buffer space By comparing the constraint lengths, if the constraint length is less than or equal to the target space, it is determined to be a safe truncation, thereby effectively reducing the false alarm rate of overflow vulnerabilities.

[0058]

[0059] (2) Completeness check of command injection purifier: Construct a predefined sanitizer rule base, and check not only whether the tainted path of command injection passes through the purifier function, but also further evaluate the completeness of the purifier, that is, whether the character blacklist filtered by the purifier function covers the complete set of shell metacharacters. (e.g., "|", ";", "&", "`", "$()", etc.). If the air purifier... filter set Incomplete coverage If this path is not found, it is still considered to be at risk of being bypassed and injected.

[0060]

[0061] Step 6: Mirror device configuration injection

[0062] Traditional text-editing attacks often fail due to the inability to obtain the decryption key, which is a common feature of modern IoT devices, in response to the configuration file encryption protection mechanisms used in their production. This invention proposes a mirror device configuration injection method that leverages the hardware and software isomorphism between devices of the same model to inject configuration data into a local verifier device. As the victim equipment The semantic converter enables transparent modification of encrypted configuration files without requiring reverse encryption algorithms or key extraction.

[0063] This method relies on a common implementation flaw: the encryption process is not bound to a unique device identifier. That is, for a group of devices of the same model, manufacturers often pre-install the same static encryption key in all devices to simplify production and maintenance. This means that the encryption function Enc(·) and decryption function Dec(·) of the configuration file are completely universal across different devices, independent of the device's unique hardware identifier. Formally, the configuration of the compromised device... With the validator's device configuration ,exist

[0064]

[0065] and exist and Shared among them. Based on this characteristic, the validator first... Obtain encrypted configuration files from the legitimate configuration export interface. ;Will Import And trigger its local configuration recovery process, making by Complete decryption and load the configuration into non-volatile memory; then, use the debug interface or firmware emulation to... The exploitable potential configuration attributes output from step 5 are written into the attack payload; finally, the call is made. A legitimate configuration export interface is compressed, hashed, and encrypted by the device itself, resulting in a malicious configuration file that is identical to the original format. and through The configuration recovery interface is uploaded to the victim device. Since the encryption algorithm and key are universal across devices of the same model, the victim device can successfully pass the integrity check and complete the decryption, thus completing the tampering with its non-volatile configuration storage.

[0066] The above six steps constitute a complete end-to-end attack chain: Steps 1-3 complete the unsupervised localization of configuration processing functions, solving the obstacle of configuration processing function identification; Steps 4-5 complete the precise screening of exploitable potential configuration attributes, solving the problem of ambiguity in potential attribute verification; Step 6 completes the transparent injection of encrypted configuration files, solving the encryption injection barrier.

[0067] Example

[0068] This invention provides an automated vulnerability discovery method for potential configuration attributes of IoT devices. This method can systematically discover potential configuration attribute vulnerabilities that are invisible to the front end but actually processed in the back end of IoT firmware across different vendors and architectures. Furthermore, it can bypass encrypted configuration protection to achieve practical-level injection. The overall process of this invention is as follows: Figure 2 As shown, the process includes three sequential stages: the location stage, the identification stage, and the utilization stage, corresponding to steps 1-3, 4-5, and 6, respectively. This embodiment targets typical gateway devices such as home routers, network attached storage (NAS), and network cameras. The specific implementation steps are as follows:

[0069] Step 1: Candidate Domain Pruning

[0070] This embodiment uses the Binwalk automation tool to unpack the target firmware image and extract the root file system. Within the root file system, this embodiment analyzes the startup instructions of typical daemons such as the system initialization scripts / etc / init.d / rcS and / etc / inittab, as well as the web service httpd, to locate the boundary binary set. .

[0071] Subsequently, this embodiment uses the ELF dependency viewing tool readelf and the binary analysis tool objdump to analyze each... Extract all linked dynamic link libraries; based on the standard library whitelist, such as libc, libm, libpthread, libdl, libssl, libcrypto, and other common basic libraries, remove non-business libraries; the remaining private shared libraries constitute the candidate domains for function behavior genealogy analysis. Original boundary binary This will be retained as the entry procedure set for the subsequent step 5, taint analysis.

[0072] Step 2: Constructing the genealogical diagram of function behavior

[0073] This embodiment uses IDA Pro to perform decompilation, symbol recovery, and call graph construction. It extracts structural and data flow features of the FBLG based on Hex-Rays AST and IDA's built-in Data Dependency Graph (DDG). Each function f in the invention is assigned one of three types of atomic behavior tags: Set-Tag, Get-Tag, or Null-Tag, according to the method described in this invention.

[0074] After completing the atomic configuration interface marking, this embodiment further introduces complexity gradient constraints and data flow convergence constraints to perform spectral convergence expansion on the candidate encapsulation functions. Specifically, as follows... Figure 3 As shown, the `nvram_get` function has a cyclomatic complexity of 18, which is typical of lightweight low-level configuration access primitives. Tracing upwards along the call graph, it can be found that functions such as `acosNvramConfig_read` (complexity 9) and `acosNvramConfig_get` (complexity 9) directly call this atomic function, and their cyclomatic complexities are significantly lower than those of the core nodes and remain within the same low-complexity gradient range. This indicates that these functions only perform parameter encapsulation, type conversion, return value normalization, or lightweight semantic adaptation outside the atomic interface, without introducing large-scale business control logic. Therefore, they are all included in the same configuration read interface family.

[0075] Furthermore, above acosNvramConfig_read and acosNvramConfig_get, there are secondary encapsulation nodes such as acosNvramConfig_get_decodecode (complexity 7). These nodes not only still satisfy the low-complexity gradient continuation feature, but their output data also continue to converge stably to the same configuration value return path, indicating that they essentially still belong to the continuous configuration semantic wrapping layer built around nvram_get. Therefore, this embodiment continues to absorb them into the configuration reading spectrum.

[0076] Through the above process, the discrete functions in the candidate private library are organized into several FBLG family trees rooted in the core parent function, which serve as the basic units for subsequent embedding and clustering.

[0077] Step 3: Locating the configuration processing function based on graph isomorphic networks

[0078] This embodiment of Random-GIN is implemented based on PyTorch-Geometric, with an embedding dimension d=128, a network layer number K=3, and a fixed random seed of 42 to ensure comparability of embeddings between different firmware. For each FBLG node, its CFG topology features, DDG data flow features, and system API behavior distribution are concatenated to form an initial feature vector, which is then propagated through a three-layer structure perception using Random-GIN.

[0079] Then, the node affinity weights are calculated according to formula (2). The full-image features are obtained by weighted aggregation according to formula (3). Within the embedding space, the OPTICS clustering algorithm provided by scikit-learn is used to identify clusters of structurally homologous functions; finally, the centrality index is used according to formula (4). Sort the data and identify the Top-1 dual family as the target. , ).

[0080] Step 4: Volatility Attribute Pruning

[0081] This embodiment uses an SDG-based engine to elevate the binary data to a normalized intermediate representation, constructing a system dependency graph. All call points located in the firmware in step 3 constitute a sink point set. For each call point, verification is performed according to the cascading order described in this invention: "L1 intra-process post-domination → L2 inter-process reachability → immutable data source constraints".

[0082] Taking the Linksys WRT54G initialization program (rc) as an example, although the function sub_4072DC in the rc file reads the charset attribute, it deterministically calls the nvram_set function at the end to write this attribute. Therefore, this embodiment first constructs the control flow graph of sub_4072DC and verifies that the nvram_set call point is a strictly dominant node after the function entry, satisfying L1 constraints. Then, tracing back the call chain, it is found that in the entire chain from the main function to sub_4072DC, each call point satisfies an L1 dominant relationship in its parent function, satisfying L2 constraints. Finally, by slicing backwards, it is confirmed that the written variable can only be taken from a hard-coded constant in the read-only data segment, i.e., one of two fixed strings, satisfying the immutable data source constraint. All three constraints are satisfied simultaneously, therefore this attribute is determined to be a volatile attribute and is removed.

[0083] Step 5: Lightweight taint analysis for configuration attributes

[0084] The lightweight taint engine in this embodiment uses the CIF_get* output parameters and return values ​​from step 3 as taint sources, and dangerous library functions and command execution functions as taint aggregation points. The engine employs fixed-point iterative propagation in the order of basic blocks within the process and is equipped with a multi-layered false alarm filtering chain, including integer conversion tracing, format string analysis, bounded copy detection, length guard identification, purifier completeness check, assignment chain tracing, constant folding, and alias approximation.

[0085] This example illustrates the workflow using a command injection vulnerability in the Western Digital My Cloud OS 5 system. Within a function of a CGI program, two network attributes are first read from an XML configuration database using CIF_get* functions. These attributes originate from an internal configuration file, not from front-end user input. The taint engine uses these two read attributes as taint sources and performs propagation analysis along the process path.

[0086] In the main processing path, these two taint attributes are merged with another user input field from an HTTP form, and the whole thing is processed by a dedicated command parameter escaping function. According to the purifier completeness judgment rules, the purifier on this path is judged to be complete, so no alarm is generated.

[0087] However, in the error handling path, these two attributes from the XML configuration database were directly concatenated into the command string and passed to the system command execution function without any cleanup. According to the cleanup purifier completeness judgment rules, no complete cleanup purifier exists on this path, so the engine identifies it as an exploitable command injection path and generates an alert. This case demonstrates that the same taint source may have different security attributes in different program paths, and path-based cleanup purifier completeness checks can effectively identify vulnerabilities caused by path branches.

[0088] Step 6: Mirror device configuration injection

[0089] First, the verifier passes. The legitimate "configuration backup / export" function downloads the encrypted configuration file. ;Will Import and through The "Configuration Restore / Import" function triggers its local configuration loading process, which is... Internally, static keys shared by the vendor are used. Complete decryption, compression and restoration, and XML / binary structure deserialization, and write the plaintext configuration according to formula (8). Non-volatile storage; afterwards, the verifier uses the UART debug interface, serial console, or firmware emulation to... The above precisely writes the exploitable potential configuration attributes output from step 5, such as writing the DNS2 field into the command injection payload; after writing, it calls... The "Configuration Backup / Export" function is provided by The serialization, compression, HMAC calculation, and AES / DES encryption are performed using the original manufacturer's logic to obtain a legitimate encrypted configuration file that is completely consistent with the original format. The final verifier will pass The "Configuration Recovery / Import" function uploads the configuration to the affected device. Due to... It is compatible with devices of the same model, and the encryption process is not bound to the device's unique identifier. Able to pass smoothly The verifier performs HMAC integrity verification and completes AES / DES decryption, thereby tampering with the non-volatile configuration storage of the victim device. The verifier then sends a legitimate configuration activation request to the victim device, which triggers the injection path identified in step 5, allowing arbitrary commands to be executed with root privileges.

[0090] In summary, this embodiment verifies that the method described in this invention can automatically mine and persistently inject potential configuration attributes end-to-end in real IoT firmware with cross-vendor, cross-architecture, and encryption protection, and has good low false alarm rate and scalability, providing a practical and feasible technical means for security hardening and proactive vulnerability discovery of IoT device groups.

Claims

1. An automated vulnerability discovery method targeting the potential configuration attributes of Internet of Things (IoT) devices, characterized in that, Includes the following steps: Step 1: Pruning candidate domains: Unpack the target firmware, identify the boundary binary related to system initialization and network services, and extract the dynamic link libraries that the binary files depend on; exclude standard library components through a whitelist mechanism, and narrow the analysis scope to the vendor's private shared libraries that carry core business logic, as candidate domains for locating configuration processing functions; Step 2: Function Behavior Genealogy Construction: Perform lightweight data flow analysis on each function in the candidate private library and assign atomic behavior labels according to the data flow of parameters and return values; apply complexity gradient constraints and data flow convergence constraints on the call relationship, and aggregate semantically related encapsulated functions and core parent functions into a function behavior genealogy. Step 3: Configuration processing function localization based on graph isomorphic network: Vectorize and embed each node in the function behavior spectrum graph using a graph isomorphic network with random parameters, and aggregate the full graph vectors using the inverse of the topological distance from the node to the parent function as the affinity weight; In the embedding space, the OPTICS clustering algorithm is used to identify nodes with structural homology and filter outliers as candidate clusters; Based on the centrality index, the candidate clusters are ranked by their in-degree and reference frequency in the global call graph, and the dual families with the highest ranking are determined as the configuration write function family and the configuration read function family; Step 4: Volatile Attribute Pruning: Promote the binary code to a normalized intermediate representation and build a system dependency graph. Treat all API call points where configurations are written as pollution points. Perform strict post-domination analysis within the process and reachability analysis between processes on each pollution point. Then, add immutable data source constraints and remove volatile attributes that are deterministically reset on the startup path. Step 5: Lightweight taint analysis for configuration attributes: Using the configuration read function family as taint sources and dangerous library functions and command execution functions as taint functions, record copy length constraints during propagation and compare them with the target buffer space to determine the reachability of buffer overflow; It performs purifier integrity checks on command injection paths based on a predefined purifier rule base; finally, it filters unusable attributes in the project using the system's key field knowledge base and outputs the final list of usable potential configuration attributes. Step 6: Mirror Device Configuration Injection: Obtain a validator local device of the same model as the victim device, import the encrypted configuration file exported by the victim device and decrypt it; modify the exploitable potential configuration attributes in the non-volatile storage of the validator local device into the attack payload; call the built-in export interface of the validator local device to re-encapsulate and encrypt it to obtain a malicious configuration file with the same encryption format as the victim device, and then upload it to the victim device to complete the injection.

2. The automated vulnerability discovery method for potential configuration attributes of IoT devices according to claim 1, characterized in that, The specific method for constructing the function behavior genealogy in step 2 is as follows: For each function in the candidate private library, perform lightweight data flow analysis on the parameters to return value and the parameters to memory write. The function whose parameter data flow is finally merged into persistent storage or memory write operation is marked as Set-Tag, the function whose return value or output parameter comes from storage or global state is marked as Get-Tag, and the rest of the functions are marked as Null-Tag. Two types of constraints are introduced to ensure semantic consistency between the encapsulated functions and the core functions: Complexity gradient constraint: requires that the cyclomatic complexity of the sub-function be strictly less than the cyclomatic complexity of the parent function; Data flow convergence constraint: requires that there be a clear bidirectional data dependency path between the wrapper function and the parent function: the return value of the wrapper function comes from the return value of the parent function, and the parameters of the parent function come from the wrapper function; To address the semantic copying scenario caused by compiler tail call optimization, the encapsulated function is treated as a semantic alias of the parent function, becoming an independent node but inheriting all feature vectors of the parent function; After the above steps, functions with the same atomic behavior tags in the candidate private library are organized into a function family rooted in the core parent function. The function family is used as the basic node unit of the function behavior genealogy graph, thus completing the construction of the function behavior genealogy graph.

3. The automated vulnerability discovery method for potential configuration attributes of IoT devices according to claim 1, characterized in that, The specific method for vectorizing and embedding each node in the function behavior spectrum graph using a graph isomorphic network with randomized parameters, and aggregating the full graph vectors using the reciprocal of the topological distance from a node to its parent function as a affinity weight, is as follows: A three-layer graph isomorphic network with randomized parameters is used to vectorize and embed each node in the function behavior spectrum graph to obtain feature vectors. ; The weighted aggregation of the full-graph eigenvectors of the function behavior spectrum is performed using the following formula: ; In the formula, The weight of the i-th node is as follows: ; Let be the topological distance from the i-th sub-function to the parent function.

4. The automated vulnerability discovery method for potential configuration attributes of IoT devices according to claim 1, characterized in that, The method for determining the highest-ranking dual family as the configuration write function family and configuration read function family based on the centrality index to sort the candidate families by in-degree and reference frequency in the global call graph is as follows: By imposing semantic completeness constraints on each candidate cluster, potential candidate dual families are obtained. That is, a potential candidate dual family is a candidate cluster that simultaneously contains the Set-Tag function family and the Get-Tag function family. Set-Tag refers to a function whose parameter data stream is ultimately merged into persistent storage or memory write operations, and Get-Tag refers to a function whose return value or output parameter comes from storage or global state. All candidate pairs are ranked using a centrality metric, and the candidate pair with the highest ranking is selected as the target pair. , The specific formula is: ; In the formula, To configure the write function family, For the configuration read function family, s represents the candidate configuration write function family, i.e., the Set-Tag function family; g represents the candidate configuration read function family, i.e., the Get-Tag function family; Cands is the set of all candidate dual families, C in (s,g) represents the sum of the in-degree and reference frequency of this pair of function families in the global call graph.

5. The automated vulnerability discovery method for potential configuration attributes of IoT devices according to claim 1, characterized in that, The system dependency graph is represented as SDG = (V, E), where V represents basic blocks and function call points, and E integrates control flow, data flow, and call graph information.

6. The automated vulnerability discovery method for potential configuration attributes of IoT devices according to claim 1, characterized in that, For each contamination point, perform strict intra-process post-domination analysis and inter-process reachability analysis sequentially, then add immutable data source constraints, and eliminate volatile attributes that are deterministically reset on the startup path, including: In-process strict post-domination analysis: Construct a control flow graph for a function f that contains contamination point A. Given a post-dominant tree, determine whether contaminated point A is a strictly post-dominant node of the entry node. This proves that as long as function f is executed, operation s will necessarily be triggered. Inter-procedural reachability analysis: Backtrack a call chain from the entry point to function f at the call graph level and verify whether each call point in the call chain satisfies the strict post-domination relationship within the procedure in the parent function. If it does, then the inter-procedural reachability is satisfied. Immutable data source constraint: Perform backward slicing on the write parameters at pollution point A, with the slice terminating at a static constant in the data segment; Attributes that simultaneously satisfy the above constraints of strict post-domination within the process, reachability between processes, and immutable data source are judged as volatile attributes and are removed.

7. The automated vulnerability discovery method for potential configuration attributes of IoT devices according to claim 1, characterized in that, The specific method for lightweight taint analysis oriented towards configuration attributes is as follows: Configure the family of read functions The output parameters and return values ​​are taint sources. Fixed-point iterative propagation is performed along the basic block sequence within the process. Each iteration records the length constraint introduced by the data transmission operation. ; When the tainted variable reaches the copy function, the most stringent cumulative constraint is compared. With target buffer space ,like If the tainted variable reaches the command execution function, the purification function on the propagation chain is traced in reverse, and it is determined whether the character blacklist covers the complete set of dangerous characters based on the predefined purifier rule base. If the filter set of the purifier does not completely cover the set of dangerous characters, the path is still determined to be at risk of being bypassed and injected.

8. The automated vulnerability discovery method for potential configuration attributes of IoT devices according to claim 1, characterized in that, The specific method for injecting mirror device configuration is as follows: Selection of the victim's equipment Local verifier devices of the same model with isomorphic hardware and software As a local semantic converter; via the victim device Obtain encrypted configuration files from the legitimate configuration export interface. ; Encrypt configuration file Import validator's local device And trigger the local configuration recovery process of the validator's local device, making the validator's local device Decryption is completed using a vendor-shared static key, and the configuration is loaded into non-volatile storage. On the local device via debug interface or firmware emulation. The above can be used to write attack payloads using potential configuration attributes; Call the validator's local device A legitimate configuration export interface is compressed, hashed, and encrypted by the device itself, resulting in a malicious configuration file that is identical to the original format. ; Finally, the malicious configuration file was... Through the victim's equipment The configuration recovery interface is uploaded by the victim device. Integrity verification and decryption are completed using the same set of shared keys, thereby enabling the tampering of the non-volatile configuration storage of the compromised device.

Citation Information

Patent Citations

  • Internet of Things firmware vulnerability mining method and device and electronic equipment

    CN116541357A

  • Router firmware two-stage vulnerability automatic mining method based on GNN and angr

    CN118797658A