Application program analysis method and device based on multi-model cooperation, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610001969.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-04
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-01-04
AI Technical Summary
而对于普通用户来说,随着身份信息等重要信息的泄露,有可能导致其财产遭受到不必要的损失
[0035]可选的,所述存储介质配置有工作区、RPA工具、网络抓包工具和反编译工具。
Smart Images

Figure CN121859312B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cybersecurity technology, and more specifically, to an application analysis method, apparatus, electronic device, and storage medium based on multi-model collaboration. Background Technology
[0002] With the development of mobile devices, applications running on these devices have brought increasing convenience, enabling people to complete many tasks. However, this convenience also provides opportunities for malicious individuals or organizations. For example, some applications maliciously collect users' identity information without their consent, facilitating the misuse of this information for illicit purposes. For ordinary users, the leakage of important information such as identity details could lead to unnecessary financial losses. The inventors of this application have discovered in practice that effective analysis of certain malicious applications can yield substantial clues, which can help users mitigate potential losses in a timely manner. Summary of the Invention
[0003] In view of this, this application provides an application analysis method based on multi-model collaboration, applied to electronic devices, the application analysis method comprising the following steps:
[0004] Receive the installation package of the application to be analyzed and place the installation package into the workspace;
[0005] The source code data of the application to be analyzed is input into the large model through the multi-hop principle, so that the large model can perform inference analysis on the source code data;
[0006] During the reasoning and analysis of the source code data by the large model, dynamic and static analysis are performed on the source code data to obtain static analysis summaries and dynamic analysis summaries. Semantic-aware RoPE dynamic scaling processing is then performed to enable the large model to obtain key information from the source code data.
[0007] Optionally, during the reasoning and analysis of the source code data by the large model, dynamic and static analysis are performed on the source code data to obtain static and dynamic analysis summaries, and semantically aware RoPE dynamic scaling processing is performed to enable the large model to obtain key information in the source code data. This includes the following multiple iterative execution steps:
[0008] The source code data is obtained by calling a decompilation tool, and the source code is analyzed using the first major model based on code symbol vectors to obtain the static analysis summary;
[0009] In a cloud phone environment, the second major model with multimodal capabilities is used to call RPA tools to install and operate applications, capture the cloud phone's running data, and analyze the running data to obtain the dynamic analysis summary;
[0010] The third model, which has the ability to invoke tools, is used to comprehensively evaluate the static analysis summary and the dynamic analysis summary in order to identify information gaps and plan the information acquisition strategy for the next round of reasoning. Then, the process returns to the initial step. When the third model determines that the currently acquired combined information is sufficient to generate the final response, the iteration is terminated and the key information is generated.
[0011] Specifically, when the third major model is used in the comprehensive evaluation, it employs semantically aware RoPE dynamic scaling technology to process combined information in order to support accurate understanding of long contexts.
[0012] Optionally, the step of obtaining the source code data by calling a decompilation tool and analyzing the source code using a first major model based on code symbol vectors to obtain the static analysis summary includes the following steps:
[0013] Obtain a list of all components in the source code data;
[0014] Based on the component list, the source code data is analyzed and processed to identify vulnerabilities.
[0015] Optionally, the step of analyzing and processing the source code data based on the component list to obtain the vulnerabilities includes the following steps:
[0016] Based on the component list, obtain a list of component class names for all components;
[0017] Based on the list of component class names, the vulnerability database is searched to obtain the corresponding vulnerability details;
[0018] Find the detailed source code corresponding to the vulnerability details in the vulnerability database;
[0019] The source code data is compared with the detailed source code to determine whether any vulnerabilities exist.
[0020] Optionally, the step of analyzing and processing the source code data based on the component list to obtain the vulnerabilities further includes the following steps:
[0021] Before comparing the source code data with the detailed source code, the source code data is processed based on the octet symbol of the additional structured vector.
[0022] Optionally, the structured vector includes a magic number part, a type part, a namespace part, and an index constant part.
[0023] Optionally, the analysis method further includes the step of:
[0024] The multiple large models are trained and enhanced using a professional vulnerability database. The multiple large models include some or all of the first large model, the second large model, and the third large model.
[0025] An application analysis device for use in an electronic device, the application analysis device comprising:
[0026] The program receiving module is configured to receive the installation package of the application to be analyzed and place the installation package into the work area;
[0027] The reasoning and analysis module is configured to input the source code data of the application to be analyzed into the large model through the multi-hop principle, so that the large model can perform reasoning and analysis on the source code data;
[0028] The dynamic scaling module is configured to perform dynamic and static analysis on the source code data during the reasoning and analysis process of the large model, obtain static analysis summary and dynamic analysis summary, and perform semantically aware RoPE dynamic scaling processing so that the large model can obtain key information in the source code data.
[0029] Optionally, the application analysis device further includes:
[0030] The model optimization module is configured to train and enhance multiple large models using a professional vulnerability database, wherein the multiple large models include some or all of the first large model, the second large model, and the third large model.
[0031] An electronic device, which is a cloud phone or a graphics card, includes at least one processor and a memory connected to the processor, wherein:
[0032] The memory is used to store computer programs or instructions;
[0033] The processor is used to execute the computer program or instructions to enable the electronic device to implement the application analysis method as described above.
[0034] A storage medium for use in an electronic device, the storage medium carrying one or more computer programs that can be executed by the electronic device, thereby enabling the electronic device to implement the application analysis method as described above.
[0035] Optionally, the storage medium is configured with a workspace, RPA tools, network packet capture tools, and decompilation tools.
[0036] As can be seen from the above technical solution, this application discloses an application analysis method, apparatus, electronic device, and storage medium based on multi-model collaboration. This method and apparatus are applied to electronic devices, specifically by receiving the installation package of the application to be analyzed and placing the package into the working area. The source code data of the application to be analyzed is input into a large model through a multi-hop principle, enabling the large model to perform inference analysis on the source code data. During the inference analysis process, the source code data undergoes dynamic and static analysis to obtain static and dynamic analysis summaries, and semantically aware RoPE dynamic scaling processing is performed to allow the large model to obtain key information from the source code data. Through these operations, applications on mobile devices can be effectively analyzed, and the obtained clues can be used to mitigate certain losses with the assistance of authorized agencies. Attached Figure Description
[0037] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1 A diagram illustrating a set of code data used to construct the body of a POST request;
[0039] Figure 2 This is a flowchart illustrating an application analysis method based on multi-model collaboration, according to an embodiment of this application.
[0040] Figure 3 This is a schematic diagram of the overall architecture of the visualized intelligent agent workflow in an embodiment of this application;
[0041] Figure 4 This is a schematic diagram illustrating the effect of implementing a large model according to an embodiment of this application;
[0042] Figure 5 This is a flowchart illustrating another application analysis method based on multi-model collaboration, as described in this application.
[0043] Figure 6 This is a block diagram of an application analysis device based on multi-model collaboration, according to an embodiment of this application.
[0044] Figure 7 This is a block diagram of another application analysis device based on multi-model collaboration according to an embodiment of this application;
[0045] Figure 8 This is a block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0046] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0047] Based on the technical requirements of this application, this section compares the actual performance of the most prominent large language models in APP analysis for decompiled code: GPT-4 performs reasonably well in source code understanding, but is limited to literal semantics and lacks deep reasoning capabilities. In other words, GPT-4 can understand code well when the code style is standardized, variable names are clear, and comments are complete; however, its performance drops significantly when faced with obfuscated, uncommented decompiled code. Figure 1 As shown in the image, the code is intended to construct a POST request body, but after obfuscation, GPT-4 can no longer recognize its true function.
[0048] To address this issue, an effective solution is to provide GPT-4 with the code snippet calling the aforementioned function, along with the relevant code from its dependent libraries (such as okhttp). By supplementing this with more comprehensive contextual information, the model can more accurately understand the code's functionality. However, this approach faces two key technical challenges.
[0049] On the one hand, there is the bottleneck of large-scale code semantic parsing: complete mobile applications are usually huge, and the decompiled code may exceed hundreds of MB, far exceeding the existing large model context window (usually millions of tokens), and will significantly increase the consumption of video memory; on the other hand, there is the decoupling of symbol semantics: obfuscated class names, variable names, and function names are often renamed to highly repetitive identifiers such as a, b, c, etc., leading to namespace pollution, breaking the context connection, and ultimately affecting the model's correct understanding of the code. Based on this, the following specific technical solutions are proposed.
[0050] Figure 2 This is a flowchart illustrating an application analysis method based on multi-model collaboration, as described in an embodiment of this application.
[0051] like Figure 2 As shown, the analysis method provided in this embodiment is applied to electronic devices to analyze and process application source code data based on the collaboration of multiple large models, thereby obtaining the necessary clues. The electronic device can be understood as a computer, server, or cloud platform with data computing and information processing capabilities. The analysis method specifically includes the following steps:
[0052] S1. Receive the installation package of the application to be analyzed from the user input and place the installation package into the workspace.
[0053] The workspace here can be seen as a temporary storage area for user configurations, such as memory or a database, so that the system can read the installation package from it to perform subsequent operations.
[0054] S2. Guide the large model analysis application through multiple principles to identify and obtain the key information needed to complete the task.
[0055] Traditional large-scale models, such as the GPT-4 language model, typically employ single-hop inference when analyzing applications (APPs), inputting all the content to be analyzed into the model at once. However, APK files usually exceed 100MB, equivalent to approximately 100M tokens, while the context window of mainstream large-scale models can only hold a maximum of about 1M tokens, and even as little as 32K tokens for high-precision applications. Efficiently analyzing 100M-level applications within a limited 32K context is a key challenge.
[0056] Based on this, this application inputs source code data using a multi-hop principle. Multi-hop reasoning extracts the most critical contextual information step by step through a "thinking and data acquisition" approach. In each step, the large language model system supporting multi-hop reasoning prioritizes acquiring the most important data context segment at the moment, and then performs in-depth analysis and reasoning within that small scope. This achieves efficient and accurate understanding of large code structures.
[0057] If represented using a visual agent workflow (such as Dify ChatFlow), its overall architecture is as follows: Figure 3 As shown. The characteristic of this workflow is that each step only obtains the most urgently needed information. For example, when the goal is to check for vulnerabilities in a 100MB APK, the traditional method requires inputting all content at once, approximately 100MB of tokens. Therefore, this application adopts a step-by-step, multi-round dialogue approach, with the specific process as follows:
[0058] Step 1: Obtain a list of all components in the source code data.
[0059] Generally, LLMs don't know the contents of the application being analyzed. Therefore, for the task of "checking for vulnerabilities in an application," the most urgently needed information is a list of the application's components. Thus, the LLM is driven to call a tool to "obtain a list of all components within the APK package," obtaining approximately 1KB of tokens; that is, first, it retrieves a list of all components.
[0060] Step 2: Obtain a list of component class names for all components based on the component list.
[0061] By analyzing the list of all components within the APK package, if several components with previously high-risk vulnerabilities are identified, the following operations are performed iteratively for each vulnerable component. Specifically, a unique key is obtained to query the vulnerability database so that the LLM can perform targeted analysis. Since our vulnerability database uses class names to index vulnerability information, the LLM's most pressing need at this point is to obtain the list of component class names (approximately 1K tokens).
[0062] Step 3: Search the vulnerability database based on the list of component class names to obtain the corresponding vulnerability details.
[0063] After obtaining the list of component class names, the LLM can use these class names to search for the corresponding vulnerability list in the vulnerability database to obtain a detailed description of the vulnerability. Then, it can use the key as an index to query the detailed vulnerability information. At this point, approximately 2KB of context is obtained.
[0064] Step 4: Locate the detailed source code corresponding to the vulnerability details in the vulnerability database.
[0065] The vulnerability database describes how to confirm the existence of a vulnerability by examining the source code. At this point, the most pressing information is the detailed source code of the class described in the vulnerability database, approximately 5KB of context. Therefore, given the detailed vulnerability information, the corresponding detailed source code is determined based on that information.
[0066] The purpose of obtaining detailed source code is to compare and analyze it with the source code data of the APK. However, the APK source code obtained here is obfuscated, and it differs significantly from the code snippets in the previous vulnerability details. Obfuscated code contains numerous repetitive, similar, or meaningless symbols, such as aaa, aab, and aac. Traditional large language models treat these symbols as multiple tokens (e.g., representing aaa as [a, a, a]) and may incorrectly classify them as similar to other symbols (such as aab and aac), leading to misunderstandings. Furthermore, some symbols may share names with code keywords (such as if and of), further increasing the difficulty of understanding. Even worse, deliberately naming symbols with names unrelated to their actual meaning can mislead the model, thus reducing the accuracy of the analysis.
[0067] To address this issue, this application designs a dedicated placeholder token. The primary function of this token is to resolve the semantic similarity of variable names at the English level, thereby enabling the LLM to accurately understand the code. Furthermore, using a single token instead of multiple tokens to store the semantic information of a single variable at the code level better maintains semantic consistency. To distinguish different tokens, this application attaches a structured vector to it, possessing the following characteristics:
[0068] (1) Structured Dimension Partitioning: The vector dimension is divided into four parts:
[0069] ,in:
[0070] The magic number part is used to declare that the vector is a structured symbolic vector; The type part specifies the type of the symbol; The namespace part records the type, namespace, and index of symbols within the current scope. The model accesses this table through an attention mechanism to avoid confusion between symbols with the same name in different scopes. The index / constant part represents the unique identifier or constant value of the symbol in the program.
[0071]
[0072] During training or inference, the embedding process can be written as:
[0073]
[0074] in: For the t-th input token, For the final embedding representation, SymEmbed(s) is the symbol embedding function, and TokenEmbed(s) is the regular word embedding function.
[0075] (2) Strong robustness against large language models: Force the model to rely on the structured dimension of symbol tokens rather than the original text similarity to avoid different symbols being misidentified as the same symbol.
[0076]
[0077] That is, the types are the same (e.g., they are all variables), the namespaces are the same (e.g., they are all in com.foo.utils), and the unique index identifiers / constant values are the same.
[0078] When calculating the total loss, we introduced an additional symbol consistency comparison term on top of the original language model loss. This term is used to enhance the model's ability to recognize the same symbol. If the model cannot correctly identify the consistency of the representation of the same symbol in different positions, we will penalize this behavior. To emphasize this point, this symbol consistency loss is given 10 times the weight of the original loss and participates in the optimization process of the total loss.
[0079] ,
[0080] in: =10 indicates that the penalty term for the sign consistency part is weighted by 10 times, and the final effect is as follows: Figure 4 As shown.
[0081] (3) Symbol Dynamic Mapping: Symbol dynamic mapping is a technique used to enhance the generalization ability of a model. In this mechanism, the symbol index is only used to distinguish different symbol ontologies and does not carry any fixed semantic information. To verify and enhance the model's independence from symbol indexes, we randomly permuted the symbol index numbers during training and inference. In this way, we can effectively prevent the model from overfitting to the symbol "numbers", thus ensuring that the model truly learns the structure and contextual semantics of the symbols, rather than their surface numbers.
[0082] In this way, the model can accurately distinguish different symbols and their contexts, improving its understanding and reasoning ability regarding obfuscated code, for example:
[0083] void main() {
[0084] cin << a;
[0085] cout >> a + 1;
[0086] }
[0087] After embedding, the code is represented as:
[0088] void [placeholder token][function vector 1]() {
[0089] [Placeholder token][Variable vector 2] >> [Placeholder token][Variable vector 3];
[0090] [placeholder token][variable vector 2] << [placeholder token][variable vector 3] + [placeholder token][constant vector value 1];
[0091] }
[0092] This method is applicable not only to decompiled code, but also to assembly code, Flutter, and ReactNative bytecode.
[0093] Step 5: Compare the source code data with the detailed source code to determine if any vulnerabilities exist.
[0094] Once the source code of the vulnerability is obtained, it is compared with the source code data to confirm whether the vulnerability exists, and the analysis results of each component are output.
[0095] S3. During the reasoning process, gaps in the acquired information are identified through multiple iterations, and new information acquisition actions are planned and executed to supplement the missing key information until the task requirements are met.
[0096] Specifically, the source code data is obtained by calling a decompilation tool, and the source code is analyzed using the first major model based on code symbol vectors to obtain a static analysis summary. During execution, a structured vector representation containing magic number, type, namespace and index or constant parts is generated for the source code data based on code symbols, and key information is extracted based on the structured vector representation to form a static analysis summary.
[0097] In the cloud phone environment, the second largest model with multimodal capabilities is used to call RPA tools to install and operate applications, capture screenshots and network traffic captured during cloud phone interaction, identify user interface elements, behavioral logic and background data interaction and other runtime data, and analyze the runtime data to obtain the dynamic analysis summary.
[0098] The third model, which has the ability to invoke tools, is used to comprehensively evaluate the static analysis summary and the dynamic analysis summary to identify information gaps and plan the information acquisition strategy for the next round of reasoning. Then, it returns to the initial step. When the third model determines that the currently acquired combined information is sufficient to generate the final response, the iteration terminates and key information is generated.
[0099] In the third major model, semantically aware RoPE dynamic scaling technology is used to process combined information during comprehensive evaluation, supporting accurate understanding of long contexts. Through dynamic scaling, the major model can identify key information in the analysis of application source code data.
[0100] In the previous analysis of the source code data, analyzing a single vulnerability required one-third of the 32K context space. Analyzing more vulnerabilities or extracting longer code snippets could lead to insufficient context space. To address this issue, the semantically aware RoPE dynamic scaling technology is introduced.
[0101] Before introducing this technology, let's briefly explain the basic principle of RoPE scaling. Currently, most mainstream large-scale language models use RoPE (Rotary Positional Embedding) to replace traditional sinusoidal positional encoding. RoPE is a method that directly encodes positional information into attention computation. It enables the model to perceive the order of tokens by applying rotational transformations to every two dimensions of the query and key.
[0102] RoPE uses a fixed-frequency encoding, with different frequencies used for different dimensions i. Specifically, RoPE's encoding method is as follows:
[0103] ,
[0104] Where: represents the angular frequency of the i-th dimension; represents the dimension of the model; (RoPE typically performs rotation encoding on even-numbered dimensions).
[0105] Location The embedding is achieved through the following formula:
[0106]
[0107] Indicates the token location index
[0108] RoPE is scalable, meaning that by dividing by a scaling factor, it can represent more location indices within a fixed length range.
[0109] Introducing semantically aware scaling ,in Indicates the first The scaling factor for the location of each token (e.g., 4.0 for the code region and 1.0 for the instruction region) is calculated as follows after scaling:
[0110] ,
[0111] in, Indicates the scaled position; Indicates the original token location index; This indicates the scaling ratio corresponding to the token (e.g., 1.0 for the instruction segment and 4.0 for the code segment).
[0112] Combining the two above, we obtain the dynamically scaled RoPE angular frequency input:
[0113]
[0114] In this way, large models can dynamically adjust the scaling ratio of positional encoding based on the semantic features of different regions, enabling them to capture information of key regions more accurately within a limited context space and improve the model's performance when processing long sequences.
[0115] However, scaling with RoPE may reduce the original performance of the model. To preserve the original performance of the model as much as possible, the scaling ratio should be adjusted. Set it to a dynamic function so that it uses a higher scaling factor for unimportant context information, a lower scaling factor (or no scaling) for important context information, a higher scaling factor for information that has been inferred, and a lower scaling factor for context information that is currently being inferred.
[0116] After performing multi-hop inference and RoPE extended context on the QwQ base model, inference capabilities may significantly decline, manifesting as symbol errors in structured JSON statements and chaotic function calls. Therefore, weighted learning is needed on the call data from multi-hop specialized tools. Weighted learning involves adding penalty weights for structured output and chaotic function calls to the loss function, increasing the weights by more than 10 times to completely eliminate these problems.
[0117] S4. Based on the task objective, generate independent summaries for each key piece of information that has been acquired, and combine all summaries to generate the final response for the task.
[0118] As can be seen from the above technical solution, this embodiment provides an application analysis method based on multi-model collaboration. This method is applied to electronic devices, specifically by receiving the installation package of the application to be analyzed and placing it in the workspace. The source code data of the application to be analyzed is input into a large model through a multi-hop principle, enabling the large model to perform inference analysis on the source code data. During the inference analysis process, the source code data undergoes dynamic and static analysis to obtain static and dynamic analysis summaries, and semantically aware RoPE dynamic scaling processing is performed to allow the large model to obtain key information from the source code data. Through these operations, applications on mobile devices can be effectively analyzed, and the obtained clues can be used to mitigate certain losses with the help of authorized agencies.
[0119] In addition, in one specific embodiment of this application, the following steps are also included: Figure 5 As shown.
[0120] S5. Use a professional vulnerability database to train and enhance the large model.
[0121] Current large-scale models, such as QwQ-32B, suffer from insufficient domain-specific knowledge, particularly in security vulnerability detection. While open-source models with 32-b parameter scale perform well on general tasks, their domain-specific knowledge remains lacking. For example, without querying vulnerability databases, LLM cannot obtain a list of components and all potentially vulnerable components. To address this deficiency, this application introduces specialized vulnerability databases, such as CVE (Common Vulnerabilities and Exposures), CNVD (China National Vulnerability Database), and CNNVD (China National Vulnerability Network Database), to supplement and train large-scale models with specialized knowledge.
[0122] Specifically, the training process includes the following steps:
[0123] (1) Data Collection and Preprocessing: Collect records containing vulnerability descriptions, impact scope, and remediation suggestions from databases such as CVE, CNVD, and CNNVD. Clean and format the collected data to ensure it is suitable for model training.
[0124] (2) Constructing a training corpus: Convert the preprocessed vulnerability information into a format suitable for model understanding, such as natural language description or structured data, to construct a training corpus.
[0125] (3) Model fine-tuning: The large language model is fine-tuned using the constructed training corpus, with a focus on training the model’s ability to understand vulnerability descriptions, locate vulnerabilities, and generate remediation suggestions.
[0126] (4) Evaluation and Optimization: The model's performance on vulnerability detection tasks is evaluated using standardized evaluation metrics such as accuracy, recall, and F1 score. Based on the evaluation results, the model structure and training strategy are further optimized.
[0127] The above methods can effectively enhance the capabilities of large models in the field of professional vulnerability knowledge, laying the foundation for their application in the security field.
[0128] In this embodiment, we used four tools to list APK components, list component class names, query vulnerability databases, and obtain source code by class name. However, in practical applications, this application can introduce more tools, specifically:
[0129] Application management: Installing and deleting applications;
[0130] Process management: Switch processes, exit background processes;
[0131] Source code search by string: Retrieve the code context containing which string;
[0132] Query the code that references a specific symbol: retrieve the code context that accesses a variable or function;
[0133] Code summary: Extract task-related information from long code snippets to reduce context usage;
[0134] Retrieve resource content: Retrieve resource content such as layout, string, and image;
[0135] Perform automated operations and: use automated control of the mobile phone model to use the mobile phone;
[0136] Screenshot: Capture a screenshot of the mobile application during this process;
[0137] Acquire network traffic: Use an automated mobile phone control model to use the mobile phone and acquire the mobile phone traffic during the process;
[0138] Get Logs: Retrieve phone operation logs;
[0139] Network management: Switch between 5G, Wi-Fi, and airplane mode;
[0140] SMS Management: Select SIM card to send and receive SMS messages;
[0141] Permission management: Granting or revoking permissions for applications;
[0142] Lock screen management: Standby, lock screen, power off;
[0143] Camera control: Input images to the phone's camera;
[0144] Using hooks: Run a Frida script to intercept a function call;
[0145] Query Developer Identifier: Retrieve the developer identifier of third-party components;
[0146] Internet access: Use a browser to further explore professional knowledge;
[0147] Full-text reasoning: 40 threads concurrently reason about the code of all classes, searching for specific information and traces of the developer's identity;
[0148] Database operations: MySQL CRUD operations for cross-session analysis;
[0149] A deeper understanding of the meaning of symbols: group together codes that facilitate understanding the purpose of a particular symbol for reasoning.
[0150] When the number of external functions available for LLM to call exceeds approximately five, the system may face the "function explosion" problem. Specifically, when faced with a large number of semantically similar functions, the large model struggles to accurately understand the specific meaning of each function, leading to ambiguity or call failures.
[0151] To address this issue, this application also proposes the following solutions:
[0152] (1) Tool aggregation and hierarchical design: Multiple tools with similar functions are merged into 5 major tool categories. Each major tool category is further divided into hierarchical sub-tools based on function, forming a multi-level sub-tool structure. In this way, the model can first select the major tool category when calling, and then select the specific sub-tool within it, thereby reducing the number of functions that need to be processed and reducing the risk of function explosion.
[0153] (2) Professional training for the base model: To enable the model to better grasp the usage methods of these aggregated major tool categories and the operational knowledge of the platform tools within the system, this application also provides professional training for the base model. Specific steps include:
[0154] Data collection and organization: Collect user manuals, operation examples, and tactical knowledge accumulated by predecessors for internal system tools, and organize them into structured training data.
[0155] Model fine-tuning: Fine-tune the base model using the prepared data to enhance the model's understanding of the tools.
[0156] Evaluation and optimization: The model's performance in tool invocation is evaluated through a series of test cases, and the training data and model parameters are further optimized based on the evaluation results.
[0157] (3) Reinforcement learning method: The reinforcement learning method is introduced to improve the model’s adaptability when facing call failure or abnormal situations based on user feedback.
[0158] Through the above measures, we have effectively solved the function explosion problem, improved the model's ability to call tools in complex tasks, and provided technical support for the efficient operation of the APP analysis assistant.
[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0160] Although the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous.
[0161] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0162] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer.
[0163] Figure 6 This is a block diagram of an application analysis device based on multi-model collaboration, according to an embodiment of this application.
[0164] like Figure 6As shown, the application analysis device provided in this embodiment is applied to an electronic device to analyze and process the source code data of an application based on the collaboration of multiple large models, in order to obtain the necessary clues. The electronic device can be understood as a computer, server, or cloud platform with data computing and information processing capabilities. Specifically, the analysis device includes a program receiving module 10, a reasoning analysis module 20, and a dynamic scaling module 30.
[0165] The program receiving module is used to receive the installation package of the application to be analyzed input by the user and place the installation package into the workspace.
[0166] The workspace here can be seen as a temporary storage area for user configurations, such as memory or a database, so that the system can read the installation package from it to perform subsequent operations.
[0167] The inference analysis module guides large model analysis applications through multi-hop principles to identify and acquire the key information needed to complete tasks.
[0168] Traditional large-scale models, such as the GPT-4 language model, typically employ single-hop inference when analyzing applications (APPs), inputting all the content to be analyzed into the model at once. However, APK files usually exceed 100MB, equivalent to approximately 100M tokens, while the context window of mainstream large-scale models can only hold a maximum of about 1M tokens, and even as little as 32K tokens for high-precision applications. Efficiently analyzing 100M-level applications within a limited 32K context is a key challenge.
[0169] Based on this, this application inputs source code data using a multi-hop principle. Multi-hop reasoning extracts the most critical contextual information step by step through a "thinking and data acquisition" approach. In each step, the large language model system supporting multi-hop reasoning prioritizes acquiring the most important data context segment at the moment, and then performs in-depth analysis and reasoning within that small scope. This achieves efficient and accurate understanding of large code structures.
[0170] If represented using a visual agent workflow (such as Dify ChatFlow), its overall architecture is as follows: Figure 3 As shown in the diagram. The characteristic of this workflow is that each step only retrieves the most urgently needed information. For example, when the goal is to check for vulnerabilities in a 100MB APK, the traditional method requires inputting all content at once, approximately 100MB of tokens. Therefore, this application adopts a step-by-step, multi-round dialogue approach, the specific process of which has been detailed above and will not be repeated here.
[0171] The dynamic scaling module is used during the inference process to identify gaps in the acquired information through multiple iterations, and to plan and execute new information acquisition actions to supplement the missing key information until the task requirements are met. Specifically, based on the task objective, it generates independent summaries for each acquired key piece of information, and integrates all summaries to generate the final response for the task.
[0172] Specifically, the source code data is obtained by calling a decompilation tool, and the source code is analyzed using the first major model based on code symbol vectors to obtain a static analysis summary. During execution, a structured vector representation containing magic number, type, namespace and index or constant parts is generated for the source code data based on code symbols, and key information is extracted based on the structured vector representation to form a static analysis summary.
[0173] In the cloud phone environment, the second largest model with multimodal capabilities is used to call RPA tools to install and operate applications, capture screenshots and network traffic captured during cloud phone interaction, identify user interface elements, behavioral logic and background data interaction and other runtime data, and analyze the runtime data to obtain the dynamic analysis summary.
[0174] The third model, which has the ability to invoke tools, is used to comprehensively evaluate the static analysis summary and the dynamic analysis summary to identify information gaps and plan the information acquisition strategy for the next round of reasoning. Then, it returns to the initial step. When the third model determines that the currently acquired combined information is sufficient to generate the final response, the iteration terminates and key information is generated.
[0175] In the third major model, semantically aware RoPE dynamic scaling technology is used to process combined information during comprehensive evaluation, supporting accurate understanding of long contexts. Through dynamic scaling, the major model can identify key information in the analysis of application source code data.
[0176] In the previous analysis of the source code data, analyzing a single vulnerability required one-third of the 32K context space. Analyzing more vulnerabilities or extracting longer code snippets could lead to insufficient context space. To address this issue, semantically aware RoPE dynamic scaling technology was introduced. Details have been described above and will not be repeated here.
[0177] As can be seen from the above technical solution, this embodiment provides an application analysis device based on multi-model collaboration. This device is applied to electronic devices, specifically receiving the installation package of the application to be analyzed and placing it in the working area. It then inputs the source code data of the application to be analyzed into a large model through a multi-hop principle, enabling the large model to perform inference analysis on the source code data. During the inference analysis process, the large model performs dynamic and static analysis on the source code data, obtaining static and dynamic analysis summaries, and performs semantically aware RoPE dynamic scaling processing to allow the large model to obtain key information from the source code data. Through these operations, applications on mobile devices can be effectively analyzed, and the obtained clues can be used to mitigate certain losses with the help of authorized agencies.
[0178] In addition, in one specific embodiment of this application, a model optimization module 40 is also included, such as... Figure 7 As shown.
[0179] The model optimization module is used to train and enhance large models using a professional vulnerability database.
[0180] Current large-scale models, such as QwQ-32B, suffer from insufficient domain-specific knowledge, particularly in security vulnerability detection. While open-source models with 32-b parameter scale perform well on general tasks, their domain-specific knowledge remains lacking. For example, without querying vulnerability databases, LLM cannot obtain a list of components and all potentially vulnerable components. To address this deficiency, this application introduces specialized vulnerability databases, such as CVE (Common Vulnerabilities and Exposures), CNVD (China National Vulnerability Database), and CNNVD (China National Vulnerability Network Database), to supplement and train large-scale models with specialized knowledge. The training process has been detailed above and will not be repeated here.
[0181] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0182] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0183] Figure 8 This is a block diagram of an electronic device according to an embodiment of this application.
[0184] The following is for reference. Figure 8 This document illustrates a structural diagram suitable for implementing the electronic device in the embodiments of this disclosure. The terminal device in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as graphics cards, cloud phones, mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. This electronic device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this disclosure.
[0185] The electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from an input device 806 into a random access memory (RAM) 803. The RAM also stores various programs and data required for the operation of the electronic device. The processing unit, ROM, and RAM are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0186] Typically, the following devices can be connected to the I / O interface: input devices including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although electronic devices with various devices are shown in the figures, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0187] This application also provides an embodiment of a computer-readable storage medium.
[0188] The aforementioned computer-readable storage medium is used in electronic devices and carries one or more computer programs. This storage medium is configured with a workspace, RPA tools, network packet capture tools, and decompilation tools. When the electronic device executes one or more of these computer programs, it receives the installation package of the application to be analyzed and places it in the workspace. Through a multi-hop principle, the source code data of the application to be analyzed is input into a large model, enabling the large model to perform inference analysis on the source code data. During this inference analysis, the large model performs dynamic and static analysis on the source code data, obtaining static and dynamic analysis summaries. Semantic-aware RoPE dynamic scaling processing is then applied to allow the large model to obtain key information from the source code data. Through these operations, applications on mobile devices can be effectively analyzed, and the obtained clues can be used to mitigate certain losses with the assistance of authorized agencies.
[0189] It should be noted that the computer-readable medium described above in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0190] In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0191] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0192] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0193] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0194] The technical solution provided by the present invention has been described in detail above. Specific examples have been used to illustrate the principle and implementation of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core idea of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation and application scope based on the idea of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. An application analysis method based on multi-model collaboration, applied to electronic devices, characterized in that, The application analysis method includes the following steps: Receive the installation package of the application to be analyzed and place the installation package into the workspace; The application to be analyzed is input into a large model through the multi-hop principle, so that the large model can perform inference analysis on the application. During the reasoning and analysis of the application's source code data by the large model, dynamic and static analysis are performed on the source code data to obtain and identify key information, and a comprehensive evaluation is conducted. Information gaps are reflected upon, and if there are key information gaps, subsequent steps are planned and the iteration is returned to supplement the key information required for the large model's reasoning. Finally, the response result of the task is generated. Specifically, during the reasoning and analysis of the application by the large model, dynamic and static analyses are performed on the application to obtain static and dynamic analysis summaries. A comprehensive evaluation is then conducted to identify information gaps. If critical information gaps exist, subsequent steps are planned and the iteration is returned to supplement the critical information required for the large model's reasoning. This includes the following multiple iterative execution steps: The source code data is obtained by calling a decompilation tool, and the source code data is analyzed using the first major model based on code symbol vectors to understand the key code intent or identify the additional code needed to thoroughly understand the code intent, thereby obtaining the static analysis summary. In a cloud phone environment, the second major model with multimodal capabilities is used to call RPA tools to install and interact with the application, triggering key code logic, capturing the cloud phone's running data, and analyzing the running data to understand the actual running logic of the key code, or to identify the additional interface interactions required to trigger the key code, thereby obtaining the dynamic analysis summary. The third model, which has the ability to call tools, is used to comprehensively evaluate the static analysis summary and dynamic analysis summary to identify missing information points and plan the information acquisition strategy for the next round of reasoning. Then, it returns to the step of obtaining the source code data by calling the decompilation tool. When the third model determines that the currently obtained combined information is sufficient to generate the final response, the iteration is terminated, and the key information, static analysis summary and dynamic analysis summary are used to generate the response result of the task. Specifically, when the third major model performs the comprehensive evaluation, it employs semantically aware RoPE dynamic scaling technology to process combined information in order to support accurate understanding of long contexts.
2. The application analysis method as described in claim 1, characterized in that, The process of obtaining the application's source code data by calling a decompilation tool and analyzing the source code data using a first major model based on code symbol vectors to obtain the static analysis summary includes the following steps: Parse the application and extract a list of all components contained in the application; Key code fragments are identified and located in the application, including logic code that processes specific inputs and outputs, user interface interaction logic code, and code containing specific string characteristics; the source code data is analyzed and processed based on the component list and the key code fragments to identify security vulnerability information and identity association clues in the application.
3. The application analysis method as described in claim 2, characterized in that, The analysis and processing of the source code data based on the component list to identify vulnerabilities specifically includes the following multi-dimensional vulnerability detection mechanisms: Based on the component list, obtain the component and class name list, retrieve the corresponding vulnerability details and detailed source code in the vulnerability database, and compare the source code data with the detailed source code to determine the existence of the vulnerability; Based on the characteristics of the vulnerability database, key parts are located in the source code, and the existence of the vulnerability is confirmed by analyzing the key code and related source code based on the vulnerability description. Track the entire process of application input and output data processing, analyze the vulnerabilities in each step, and search for relevant vulnerabilities in the vulnerability database to confirm their existence.
4. The application analysis method as described in claim 3, characterized in that, The analysis and processing of the source code data based on the component list and key code is a comprehensive analysis process that uses the first major model to perform deep semantic understanding and logical reasoning on the source code data. The first model employs a symbolic representation mechanism based on structured vectors. When processing the source code data, it automatically adds structured vector terms as placeholders to the code symbols. The structured vector terms define the namespace and type attributes of the symbols, so as to accurately distinguish between symbols with the same name and similar shapes during the reasoning process and eliminate the interference of symbol ambiguity on the model's reasoning.
5. The symbolic representation mechanism based on structured vectors as described in claim 4, characterized in that, The structured vector includes a magic number part, a type part, a namespace part, and an index / constant part.
6. The application analysis method according to any one of claims 1 to 5, characterized in that, Some or all of the first, second, and third major models used to execute the multi-model collaborative reasoning process are deep neural network models pre-trained and enhanced using a professional vulnerability database, thereby embedding domain-specific knowledge features targeting application security vulnerabilities.
7. An application analysis device for an electronic device, characterized in that, The application analysis device includes: The program receiving module is configured to receive the installation package of the application to be analyzed and place the installation package into the work area; The reasoning and analysis module is configured to input the source code data of the application to be analyzed into the large model through the multi-hop principle, so that the large model can perform reasoning and analysis on the source code data; The model training module is configured to perform network structure adjustment and fine-tuning adaptation operations on the large model, or to enhance the training of the large model using domain knowledge; The knowledge base module is configured to build and maintain a security analysis dataset containing the latest detailed vulnerability descriptions, vulnerability exploitation code, and a database of identity clues, providing data support for the inference analysis and training enhancement of the large model. The dynamic scaling module is configured to perform dynamic and static analysis on the source code data of the application during the reasoning and analysis process of the large model, obtain and identify key information, conduct a comprehensive evaluation, reflect on information gaps, and if there are key information gaps, plan subsequent steps and return to the iteration to supplement the key information required for the reasoning of the large model, and finally generate the response result of the task. During the reasoning and analysis of the application by the large model, dynamic and static analysis are performed on the application to obtain static and dynamic analysis summaries. A comprehensive evaluation is then conducted to identify information gaps. If critical information gaps exist, subsequent steps are planned and the iteration is repeated to supplement the critical information required for the large model's reasoning. This includes the following iterative execution steps: obtaining the source code data by calling a decompilation tool, and analyzing the source code data using a first large model based on code symbol vectors to understand the key code intent or identify the additional code needed to thoroughly understand the code intent, thus obtaining the static analysis summary; in the cloud phone environment, using a second large model with multimodal capabilities, an RPA tool is called to install and interact with the application, triggering key code logic and capturing the cloud phone's operation. The system collects and analyzes the running data to understand the actual running logic of the key code or identify the additional interface interactions required to trigger the key code, thus obtaining the dynamic analysis summary. A third model with tool-calling capabilities is then used to comprehensively evaluate the static and dynamic analysis summaries to identify missing information and plan the information acquisition strategy for the next round of reasoning. The system then returns to the step of obtaining the source code data by calling a decompilation tool. When the third model determines that the currently acquired combined information is sufficient to generate the final response, the iteration terminates, and the response result of the task is generated using the key information, the static analysis summary, and the dynamic analysis summary. During the comprehensive evaluation, the third model employs semantically aware RoPE dynamic scaling technology to process the combined information, supporting accurate understanding of long contexts.
8. The application analysis device as described in claim 7, characterized in that, When the model training module performs network structure adjustment operations on the large model, it is specifically configured as follows: The network architecture of the large model is embedded with a semantically aware RoPE dynamic scaling mechanism to adjust contextual attention, and a symbolic representation mechanism based on structured vectors to eliminate code symbol ambiguity. The network structure is then fine-tuned to maintain model performance, or vulnerability knowledge is introduced for domain-enhanced training.
9. An electronic device, characterized in that, The electronic device is a cloud phone, comprising at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs or instructions; The processor is used to execute the computer program or instructions to enable the electronic device to implement the application analysis method as described in any one of claims 1 to 6.
10. A storage medium used in an electronic device, characterized in that, The storage medium carries one or more computer programs that can be executed by the electronic device, thereby enabling the electronic device to implement the application analysis method as described in any one of claims 1 to 6.
11. The storage medium as claimed in claim 10, characterized in that, The storage medium is configured with a workspace and tool components for supporting in-depth application analysis, the tool components including: Cloud-based mobile virtualization environment: A cloud-based emulation device used to provide the runtime environment for applications; Robotic process automation tools: used to simulate user interactions with applications; Network communication data interception and tampering tools: used to capture network traffic and modify and replay data packets; Reverse engineering tools: used to restore application installation packages to source code data.
Citation Information
Patent Citations
Large model assisted static code scanning result analysis method and device, electronic equipment and computer readable storage medium
CN119690807A
Large language model multi-agent cooperative work method and system
CN120450059A