Modularized architecture-based OpenHarmony application dynamic analysis framework and method
Through the modular architecture OpenHarmony application dynamic analysis framework, the problem that existing tools cannot adapt to OpenHarmony is solved, efficient test coverage and defect detection are achieved, and the flexibility of testing strategies and the ability of multi-dimensional data collaborative verification is improved.
Patent Information
- Application Number
- CN202510375959.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
AI Technical Summary
Existing dynamic analysis tools cannot adapt to OpenHarmony's Stage model and ArkUI declarative syntax, resulting in low test coverage, insufficient defect detection accuracy, and lack of the possibility of collaborative verification of multi-dimensional data.
Design an OpenHarmony application dynamic analysis framework based on a modular architecture, including code insertion module, executor module, data processing module and event generation module, supports ArkUI component recognition, dynamic behavior modeling and policy flexibility, and build dynamic page transformation diagrams through abstract syntax tree analysis, event simulation, data collection and processing to realize multi-dimensional data collaborative verification.
It improves the test coverage and defect detection efficiency of OpenHarmony applications, provides flexible testing strategy adaptation, and realizes efficient dynamic analysis and precise defect positioning of OpenHarmony applications.
Smart Images

Figure CN120295918A_ABST
Abstract
Description
Technical Field
[0001] The present invention provides an OpenHarmony application dynamic analysis framework and method based on a modular architecture, belonging to the technical field of dynamic analysis tools. Background Art
[0002] Existing dynamic analysis tools (such as Monkey and Dynodroid for Android) achieve application behavior analysis through random event generation or model-driven testing. However, these tools rely on the Android Activity / Fragment model and Java / Kotlin language features for implementation, and cannot adapt to the Stage model and ArkTS language features of OpenHarmony. Some research attempts to construct a page transition graph (PTG) through static analysis, but lacks the ability of dynamic adjustment, resulting in low test coverage and inability to effectively capture runtime crashes.
[0003] Secondly, in the prior art, due to incompatible architecture designs, traditional dynamic analysis tools are difficult to effectively parse the unique Stage model and ArkUI declarative syntax of OpenHarmony, resulting in the failure of event simulation and page recognition functions; at the same time, the existing methods for constructing a page transition graph (PTG) mainly rely on static code analysis and cannot dynamically capture the actual runtime page jump path and distributed interaction logic, resulting in the omission of tests for deep code branches and asynchronous scenarios. In addition, a single test strategy cannot adapt to a wide range of test scenarios, resulting in poor generalization of the test strategy and lack of customization ability for specific applications. Moreover, the isolated storage and analysis methods of logs, performance metrics, and page jump data further limit the possibility of multi-dimensional data collaborative verification, making the defect location accuracy insufficient. Summary of the Invention
[0004] The present invention proposes a modular dynamic analysis framework, specifically designed for OpenHarmony applications, to solve the compatibility problems caused by architecture differences in existing tools and improve test coverage and defect detection efficiency. The specific objectives include:
[0005] 1. Compatibility with OpenHarmony features: Support ArkUI component recognition, Stage model navigation, and ArkTS language instrumentation.
[0006] 2. Dynamic behavior modeling: Construct a page transition graph (PTG), track the application state in real time, and guide event generation.
[0007] 3. Policy flexibility: Provide a policy adapter to support users to customize test policies (such as coverage priority, path exploration).
[0008] The specific technical solution is as follows:
[0009] OpenHarmony Application Dynamic Analysis Framework Based on Modular Architecture, including: Code Instrumentation Module CIM, Executor Module EM, Data Processing Module DPM, Event Generation Module EGM;
[0010] Code Instrumentation Module CIM, which is the basic module for realizing dynamic analysis, is responsible for instrumenting the ArkTS source code of OpenHarmony applications. The process includes: AST analysis and instrumentation point identification, dynamic code injection, and plug-in integration.
[0011] Executor Module EM is responsible for controlling the application to run, simulating user interactions, and collecting runtime data, and includes the following sub-modules: Event Simulator, Data Collector, and Result Generator.
[0012] Data Processing Module DPM is responsible for processing the collected data. For UI-type data, the PTG generator parses the UI hierarchy, constructs a dynamic page transition graph, and guides event generation. For text-type data, DataHub stores the runtime data in a structured manner to support policy optimization.
[0013] Event Generation Module EGM, in order to customize strategies for specific scenarios, EGM provides a policy adapter, which adopts a four-stage pipeline, namely candidate generation → filtering → scoring → selection, and supports greedy DFS and random exploration strategies. Combining the PTG and DataHub data, it dynamically adjusts the event generation logic to form a "generate-execute-analyze" closed loop.
[0014] OpenHarmony Application Dynamic Analysis Method Based on Modular Architecture, including the following steps:
[0015] S1. First, the Code Instrumentation Module CIM traverses the application source code through abstract syntax tree (AST) parsing technology to identify key code points of function declarations, branch statements, and loop structures. Each AST node is attached with a unique identifier to ensure accurate positioning of the instrumentation location.
[0016] Secondly, insert monitoring code snippets at the identified instrumentation points. The instrumented AST regenerates the executable code to ensure non-invasive collection of runtime data.
[0017] Then, encapsulate the instrumentation function into an independent plug-in BJC, which supports seamless integration with the official OpenHarmony IDE. Developers can enable instrumentation with one click in the IDE, view the coverage heat map and code execution path in real time, and improve testing efficiency.
[0018] S2. To operate on OpenHarmony applications, a series of event simulations are carried out, including UI event simulation, system event triggering, and hardware key simulation, so as to completely restore all scenarios that may occur in human-computer interaction. The event simulator of the Executor Module EMD supports click, long press, swipe, and text input operations based on screen coordinates (x, y) or component IDs, covering all interaction scenarios under the ArkUI framework. The application life cycle management is performed through the Harmony Device Connector tool to simulate device state changes. And hardware events such as volume adjustment and power key operations are generated to verify the compatibility of the application with physical keys.
[0019] When the event execution is completed, the state of the application will change accordingly. The EM needs to collect the latest state to the local through the data collector. This data collector starts an FTP service in the application sandbox to transmit the coverage data to the local server in real time, bypassing the sandbox storage limit.
[0020] In addition to the coverage data, crash logs, screenshots, and application logs will also be collected. The screenshot is obtained by executing a specific instruction to capture the current GUI, while the application log is directly obtained because it is stored in the system data area / data. The crash logs include system-level crash logs, ArkTS exception stacks, and performance metrics. Through the fusion of these multi-source data, the comprehensiveness of the data is ensured.
[0021] Finally, the result generator integrates and analyzes these data to generate a visualization report. The result generator displays the coverage report in HTML format by file / function / branch level, presents the PTG navigation path graph of the page jump logic in the form of a directed graph, and the crash details report including stack trace, trigger events, and device context.
[0022] S3. The Data Processing Module DPM performs structured processing and in-depth analysis on the original runtime data. The runtime data is divided into two major categories, UI type and text type. The PTG generator in DPM is used to process UI type data, and DataHub is used to process text type data.
[0023] For UI type data, first perform UI hierarchical parsing, extract the three-level structure of Ability, Page, and Component from the UI component tree obtained by the uttest tool, and identify the foreground page. Then construct a dynamic graph, with pages as nodes and events as edges, to construct a page transition graph, recording event parameters, timestamps, and execution results. Support the shortest path algorithm to optimize the navigation logic and avoid repeated event execution. At the same time, maintain state tracking, maintain the historical state snapshot of the PTG, support rollback and branch coverage analysis, and guide the EGM to generate high-value test cases.
[0024] For text-type data, first design a data model, use a time-series database to store logs, coverage, and performance data, and establish an index to optimize query efficiency. Secondly, provide an API for the EGM to call.
[0025] S4. The event generation module EGM realizes intelligent event generation through a policy adapter. Its core mechanism is a four-stage pipeline architecture, including candidate generation, filtering, scoring, and selection.
[0026] First, in the candidate generation stage, based on the current PTG state and DataHub historical data, enumerate all possible UI events, system events, and hardware operations.
[0027] Then, in the filtering stage, eliminate invalid events and high-risk operations.
[0028] Secondly, in the scoring stage, assign priority scores to candidate events according to coverage gain, crash detection potential, and path exploration value.
[0029] Finally, in the selection stage, use greedy DFS, random exploration, or user-defined strategies to select the optimal event for execution.
[0030] In addition, provide a dynamic feedback closed-loop that allows real-time adjustment of strategies. After each event execution, the EGM receives the updated PTG and DataHub data from the DPM, and adjusts the candidate generation rules and scoring weights in real time. And support policy plugin extension.
[0031] Design event generation strategies based on the four-stage pipeline architecture and the dynamic feedback closed-loop. Brief Description of the Drawings
[0032] Figure 1 It is the overall architecture diagram of the OpenHarmony application dynamic analysis framework of the present invention;
[0033] Figure 2 It is the diagram showing the types and positions of code coverage instrumentation of the present invention;
[0034] Figure 3 It is the schematic diagram of the FTP-based sandbox data channel of the present invention;
[0035] Figure 4 It is the flowchart of the dynamic construction of PTG of the present invention;
[0036] Figure 5 It is the greedy DFS algorithm based on the four-stage pipeline of the policy adapter of the present invention. Detailed Embodiment
[0037] The invention designs and implements an OpenHarmony application dynamic analysis framework. AsFigure 1 Shows the OpenHarmony application dynamic analysis framework based on a modular architecture, which mainly includes four major modules: the Code Instrumentation Module (CIM), the Executor Module (EM), the Date Processing Module (DPM), and the Event Generation Module (EGM).
[0038] The Code Instrumentation Module (CIM) is the basic module for the framework to achieve dynamic analysis and is responsible for code instrumentation of the ArkTS source code of OpenHarmony applications. Its core processes include: AST analysis and instrumentation point identification, dynamic code injection, and plug-in integration.
[0039] The Executor Module (EM) is responsible for controlling the application to run, simulating user interactions, and collecting runtime data, and includes the following sub-modules: event simulator, data collector, and result generator.
[0040] The Data Processing Module (DPM) is responsible for processing the collected data. For UI type data, the PTG generator parses the UI hierarchy (Ability - Page - Component), constructs a dynamic page transition graph, and guides event generation. For text type data, DataHub stores the runtime data (logs, coverage rate, performance metrics) in a structured manner to support policy optimization.
[0041] The Event Generation Module (EGM), in order to customize policies for specific scenarios, EGM provides a policy adapter, adopts a four-stage pipeline (candidate generation → filtering → scoring → selection), and supports strategies such as greedy DFS and random exploration. And combines PTG and DataHub data to dynamically adjust the event generation logic to form a "generation - execution - analysis" closed loop.
[0042] The following details the analysis steps of the above analysis framework:
[0043] S1. The Code Instrumentation Module (CIM) first traverses the application source code through abstract syntax tree (AST) parsing technology to identify key code points such as function declarations (FunctionDeclaration), branch statements (If / Switch), and loop structures (For / While). Each AST node is attached with a unique identifier to ensure accurate positioning of the instrumentation location. Figure 2 Shows different inserted points under different types of code coverage requirements.
[0044] Secondly, insert monitoring code snippets at the identified instrumentation points, such as coverage counters (e.g., InstRegion, InstBranch) or performance probes (e.g., memory / CPU monitoring logic). The instrumented AST is regenerated into executable code to ensure non-invasive runtime data collection.
[0045] Then, encapsulate the instrumentation function as an independent plugin BJC, which supports seamless integration with the official OpenHarmony IDE (DevEcoStudio). Developers can enable instrumentation with one click in the IDE and view the coverage heatmap and code execution path in real time to improve testing efficiency.
[0046] S2. Executor Module (EM). To operate on OpenHarmony applications, a series of event simulations are required, including UI event simulation, system event triggering, and hardware key simulation, so as to fully restore all scenarios that may occur in human-computer interaction. The event simulator supports click, long press, swipe, and text input operations based on screen coordinates (x, y) or component IDs, covering all interaction scenarios under the ArkUI framework. The application lifecycle management (start, switch, terminate) is executed through the Harmony Device Connector (hdc) tool, and the device state changes (such as network switch, sensor response) are simulated. Hardware events such as volume adjustment and power key operation are generated to verify the compatibility of the application with physical keys.
[0047] When the event execution is completed, the state of the application will change accordingly. EM needs to collect the latest state to the local through a data collector. Since OpenHarmony applications all run in a sandbox and the generated data is deleted when the application ends and the sandbox is destroyed, it is impossible to obtain the data generated by the application. Therefore, this data collector starts an FTP service in the application sandbox to transfer the coverage data to the local server in real time, bypassing the sandbox storage limitation, as Figure 3 shown.
[0048] In addition to the coverage data, crash logs, screenshots, application logs, etc. will also be collected. Among them, the screenshot is to obtain the current GUI by executing a specific screen capture instruction, and the application log can be directly obtained because it is stored in the system data area / data. The crash logs include system-level crash logs (POSIX signal capture), ArkTS exception stacks, and performance metrics (memory occupancy, CPU utilization). Through the fusion of these multi-source data, the comprehensiveness of the data is ensured.
[0049] Finally, the result generator integrates and analyzes this data to generate a visual report. The result generator displays the coverage report in HTML format by file / function / branch, presents the PTG navigation path diagram of the page jump logic in the form of a directed graph, and provides a crash details report including stack trace, trigger events, and device context.
[0050] S3. The Data Processing Module (DPM) performs structured processing and in-depth analysis on the original runtime data. These runtime data can be divided into two categories: UI type and text type. The PTG generator in DPM is used to process UI-type data, and DataHub is used to process text-type data.
[0051] For UI-type data, as Figure 4 shown, first, perform UI hierarchy parsing, extract the three-level structure of Ability, Page, and Component from the UI component tree obtained from the uttest tool, and identify the foreground page (based on the maximum screen area determination strategy). Then construct a dynamic graph, with pages as nodes and events as edges, to construct a page transition graph (PTG), recording event parameters, timestamps, and execution results. Support the shortest path algorithm to optimize the navigation logic and avoid repeated event execution. At the same time, maintain state tracking, maintain historical state snapshots of the PTG, support rollback and branch coverage analysis, and guide the EGM to generate high-value test cases.
[0052] For text-type data, first design a data model, use a time-series database to store logs, coverage, and performance data, and establish an index to optimize query efficiency. Secondly, provide APIs for the EGM to call, such as obtaining a list of uncovered code branches, querying high-frequency crash scenarios, and supporting the generation of adaptive test strategies.
[0053] S4. The Event Generation Module (EGM) realizes intelligent event generation through a policy adapter. Its core mechanism is a four-stage pipeline architecture, including candidate generation, filtering, scoring, and selection. First, in the candidate generation stage, based on the current PTG state and DataHub historical data, enumerate all possible UI events, system events, and hardware operations. Then, in the filtering stage, eliminate invalid events (such as non-clickable components, repeated paths) and high-risk operations (such as frequent application restarts). Secondly, in the scoring stage, assign priority scores to candidate events according to coverage gain, crash detection potential, and path exploration value. Finally, in the selection stage, use greedy DFS, random exploration, or user-defined strategies to select the optimal event for execution.
[0054] In addition, a dynamic feedback closed loop is provided to allow real-time adjustment of the strategy. After each event execution, the EGM receives the PTG updated by the DPM and the DataHub data, and adjusts the candidate generation rules and scoring weights in real time. It also supports the extension of strategy plugins, such as integrating a reinforcement learning model to optimize the scoring mechanism, or combining with an LLM to generate semantic test cases.
[0055] Based on the four-stage pipeline architecture and the dynamic feedback closed loop, an efficient event generation strategy can be designed. Taking greedy DFS as an example. This greedy DFS algorithm realizes event generation through a depth-first search framework combined with a local optimal selection strategy. The algorithm starts from the current page, traverses all interactive components to generate candidate events (such as clicks, inputs, etc.). If the current page has been fully explored, a return event (BACK_EVENT) is added to trigger the backtracking mechanism; then the executed events are filtered to avoid duplication. If the candidate set is empty, it goes back to the previous page to search again. Otherwise, the event scores are calculated according to the preset weights (such as component visibility, historical trigger frequency), and the events with the highest scores are preferentially executed after sorting the scores in descending order. This process loops until the time budget is exhausted, which not only retains the exploration ability of DFS for deep paths but also selects local optimal events at each step through the greedy strategy, balancing the exploration efficiency and the depth of path coverage, and ensuring the maximum triggering of potential defect scenarios within a limited time.
[0056] The present invention adopts the design of a strategy adapter in the event generation module, allowing users to use custom strategies. In addition to greedy DFS, it can be replaced with a reinforcement learning model or an LLM-based intelligent strategy.
[0057] The present invention builds in an FTP service for real-time collection when collecting application sandbox data. The FTP service can be replaced with WebSocket or a custom RPC protocol to adapt to different device environments.
[0058] The code instrumentation module of the present invention supports custom instrumentation points and instrumentation code, and supports instrumentation at the line, function, or component level, balancing performance and monitoring accuracy.
Claims
1. An OpenHarmony application dynamic analysis framework based on a modular architecture, including: Code Instrumentation Module CIM, Executor Module EM, Data Processing Module DPM, Event Generation Module EGM; The Code Instrumentation Module CIM is the basic module for realizing dynamic analysis and is responsible for instrumenting the ArkTS source code of OpenHarmony applications; The process includes: AST analysis and instrumentation point identification, dynamic code injection, and plug-in integration; The Executor Module EM is responsible for controlling the application running, simulating user interactions, and collecting runtime data, including the following sub-modules: Event Simulator, Data Collector, and Result Generator; The Data Processing Module DPM is responsible for processing the collected data; for UI-type data, the PTG generator parses the UI hierarchy, constructs a dynamic page transition graph, and guides event generation; for text-type data, DataHub structurally stores runtime data to support policy optimization; For the Event Generation Module EGM, to customize policies for specific scenarios, EGM provides a policy adapter that adopts a four-stage pipeline, namely candidate generation → filtering → scoring → selection, and supports greedy DFS and random exploration strategies; combined with PTG and DataHub data, it dynamically adjusts the event generation logic to form a "generate-execute-analyze" closed loop.
2. The dynamic analysis method of OpenHarmony applications based on a modular architecture is characterized in that It includes the following steps: S1. First, the Code Instrumentation Module CIM traverses the application source code through Abstract Syntax Tree (AST) parsing technology to identify key code points of function declarations, branch statements, and loop structures; each AST node is attached with a unique identifier to ensure accurate positioning of the instrumentation location; Secondly, insert monitoring code snippets at the identified instrumentation points, and the instrumented AST regenerates executable code to ensure non-invasive collection of runtime data; Then, encapsulate the instrumentation function into an independent plug-in BJC to support seamless integration with the official OpenHarmony IDE; developers can enable instrumentation with one click in the IDE, view the coverage heat map and code execution path in real time, and improve testing efficiency; S2. To operate on OpenHarmony applications, a series of event simulations are carried out, including UI event simulation, system event triggering, and hardware key simulation, so as to completely restore all scenarios that may occur in human-computer interaction; the event simulator of the Executor Module EMD supports click, long press, swipe, and text input operations based on screen coordinates (x, y) or component ID, covering all interaction scenarios under the ArkUI framework; perform application lifecycle management through the Harmony Device Connector tool to simulate device state changes; and generate hardware events such as volume adjustment and power key operations to verify the compatibility of the application with physical keys; When the event execution is completed, the state of the application will change accordingly, and EM needs to collect the latest state to the local through the data collector; this data collector starts an FTP service in the application sandbox to transmit coverage data to the local server in real time, bypassing the sandbox storage limit; Finally, the result generator integrates and analyzes these data to generate a visual report; The result generator displays the coverage report in HTML format by file / function / branch, presents the PTG navigation path graph of the page jump logic in the form of a directed graph, and the crash details report including stack trace, trigger events, and device context; S3. The Data Processing Module (DPM) performs structured processing and in-depth analysis on the original runtime data; S4. The Event Generation Module (EGM) realizes intelligent event generation through a policy adapter. Its core mechanism is a four-stage pipeline architecture, including candidate generation, filtering, scoring, and selection; Design event generation strategies based on the four-stage pipeline architecture and dynamic feedback closed-loop; 3. The dynamic analysis method for OpenHarmony applications based on a modular architecture according to claim 2, wherein In S2, in addition to coverage data, crash logs, screenshots, and application logs are also collected. Among them, screenshots are obtained by executing specific screen capture instructions to capture the current GUI, and application logs are directly obtained because they are stored in the system's data area / data; crash logs include system-level crash logs, ArkTS exception stacks, and performance metrics; through the fusion of these multi-source data, the comprehensiveness of the data is ensured.
4. The method for dynamically analyzing OpenHarmony applications based on a modular architecture according to claim 2, wherein, In S3, the runtime data is divided into two categories, UI type and text type; the PTG generator in DPM is used to process UI type data, and DataHub is used to process text type data.
5. The dynamic analysis method for OpenHarmony applications based on a modular architecture according to claim 4, wherein, For UI type data, first perform UI hierarchical parsing, extract the three-level structure of Ability, Page, and Component from the UI component tree obtained from the uttest tool, and identify the foreground page; then construct a dynamic graph, with pages as nodes and events as edges, construct a page transition graph, and record event parameters, timestamps, and execution results; Support the shortest path algorithm to optimize the navigation logic and avoid repeated event execution; at the same time, maintain state tracking, maintain historical state snapshots of the PTG, support rollback and branch coverage analysis, and guide the EGM to generate high-value test cases.
6. The dynamic analysis method for OpenHarmony applications based on a modular architecture according to claim 4, wherein For text type data, first design a data model, use a time series database to store logs, coverage, and performance data, and establish an index to optimize query efficiency; secondly, provide an API for the EGM to call.
7. The method for dynamically analyzing OpenHarmony applications based on a modular architecture according to claim 4, wherein The specific process of S4 is as follows: First, in the candidate generation stage, based on the current PTG state and DataHub historical data, enumerate all possible UI events, system events, and hardware operations; Then, in the filtering stage, eliminate invalid events and high-risk operations; Secondly, in the scoring stage, assign priority scores to candidate events according to coverage gain, crash detection potential, and path exploration value; Finally, in the selection stage, use greedy DFS, random exploration, or user-defined strategies to select the optimal event for execution; In addition, provide a dynamic feedback closed-loop to allow real-time adjustment of strategies; After each event execution, the EGM receives the updated PTG and DataHub data from the DPM, and adjusts the candidate generation rules and scoring weights in real time; and supports policy plugin extension.