Dynamic unstructured key leakage detection method in applet
By combining key behavior semantic analysis and static information flow analysis technology, the key usage semantic graph of the applet is built, which solves the problem of dynamic unstructured key leakage detection in the applet, and achieves fast and resource-efficient key leakage detection.
Patent Information
- Application Number
- CN202510254872.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to effectively detect the leakage of dynamic unstructured keys in mini-programs, especially in complex environments with multi-terminal interactions. The forms of dynamic key leakage are complex and diverse and highly concealed, and it is difficult to detect through methods such as rule matching.
Combining the semantic analysis of the behavior of mini-program keys and traditional static information flow analysis technology, a semantic graph for key usage of server and client is built, and the leakage of dynamic unstructured keys in mini-programs is detected through graph similarity analysis.
It realizes efficient detection of dynamic unstructured key leakage in mini programs, can maintain fast detection speed in large-scale batch and offline analysis, and has low resource consumption, effectively reducing the risk of key leakage.
Smart Images

Figure CN120165867A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of program information security detection, and particularly relates to a method for detecting the leakage of unstructured and dynamically generated keys in small programs. Background Art
[0002] In the new mobile ecological model of "super application + small program", the small program platform has designed a unique key mechanism for identity authentication and secure communication. This mechanism usually includes three types of keys: access keys (for service calls), encryption keys (for data communication), and root keys (for generating and managing other keys). Since this new ecology involves multi-terminal interactions among the small program client, small program server, super application client, and super application server, this multi-terminal collaboration mode exacerbates the complexity of key management. Once developers have omissions in any link of key management, it may lead to the exposure of sensitive data and even endanger the security of the entire system.
[0003] Unsafe practices of developers in key management usually lead to two types of key leakage: hard-coded key leakage and dynamic key leakage. Hard-coded key leakage means that developers directly embed the key into the code of the small program client, such as writing the key into the global configuration file. This practice makes the key easily discovered by decompilation tools or code review tools and thus exposed to attackers. Dynamic key leakage occurs during the interaction between the small program client and the server, where the key is captured during the communication process. Attackers can obtain sensitive keys by triggering specific function points and intercepting network traffic. This type of leakage is usually caused by developers using key behaviors that should be stored on the server on the client side, thus triggering dynamic leakage of cross-terminal interactions. The form of dynamic key leakage is complex and diverse and highly concealed, and it is often difficult to detect.
[0004] For hard-coded key leakage, the key itself or its context usually has specific fixed characteristics, and it can be detected by means of rule scanning such as regular expressions. However, for dynamic key leakage, since the key is usually captured during the communication process and its content does not directly appear in the code, it cannot be effectively detected by methods such as rule matching. Summary of the Invention
[0005] The purpose of the present invention is to provide a detection method for dynamically unstructured key leakage in small programs that supports large-scale batch and offline analysis, has a fast detection speed, and low resource consumption.
[0006] The method for detecting dynamically unstructured key leakage in small programs provided by the present invention, first for the convenience of explanation, summarizes the key types into three categories:
[0007] (1) Root Credential (root key), a key used to generate and manage other keys;
[0008] (2) Access Credential (access key), a key used to verify user resource access permissions;
[0009] (3) Cryptographic Credential (encryption key), a key used for data encryption.
[0010] This is because there are many types of mini-program keys, and keys of different types perform different functions.
[0011] The present invention analyzes and detects different behavioral semantics for different categories of keys.
[0012] The method for detecting dynamic unstructured key leakage in a mini-program provided by the present invention combines mini-program key behavioral semantic analysis with traditional static information flow analysis technology to detect the leakage of dynamic unstructured keys in a mini-program. Its architecture is as shown in the appendix Figure 1 As shown, the entire detection process can be divided into three stages.
[0013] (1) Server-side key behavioral semantic understanding stage, specifically including:
[0014] (1) Mini-program development document collection: First, automatically parse the mini-program development documents to obtain the semantic information contained in each server-side application programming interface (API); including the name, parameters, and function description of the API. Since the mini-program development documents present each API in well-structured HTML elements and different APIs have similar organizational structures, a fixed structural pattern can be summarized and sorted out by analyzing each developer document, and then extended to the automated analysis of other APIs, and large-scale crawling of document information can be performed. Due to the good organization of the documents, this method can be easily extended to multiple super application platforms.
[0015] (2) Mini-program development document analysis: Next, understand the server-side key usage semantics based on the API information crawled above. To facilitate mini-program developers to implement server-side operations, the mini-program development documents usually explain how to obtain the parameters in the API description. Specifically, a hyperlink pointing to the source of the parameter is usually directly attached in the parameter description of a certain API.
[0016] Therefore, based on such observations, the present invention designs a heuristic strategy to extract the data dependencies between server-side APIs. For example, when the development document describes the parameter loginCode of an API, the detailed description of loginCode contains a hyperlink that directly points to its source API getLoginCode. Through this strategy, different API operations can be interconnected to construct a complete API behavior sequence. The present invention starts from the entry API and identifies the relevant dependent server-side APIs according to its parameter description. This process is iteratively repeated until no new dependent APIs are found. After the above steps, considering that the description information in the development document may lack corresponding hyperlinks, in order to comprehensively restore the data dependencies between server-side APIs, the present invention further utilizes the naming similarity of variables to establish the data dependencies between different APIs. For example, in the WeChat platform, all APIs name the encryption key as session_key. Therefore, we can analyze the data dependencies between different APIs based on this similar naming.
[0017] (3) Construction of the key usage semantic graph: To represent the API-level semantics and consider the correlation between server-side APIs, the present invention designs a new graph, called the key usage semantic graph, abbreviated as CSG. The key usage semantic graph is used to represent the semantics of the applet server side, which contains operations related to keys. The present invention defines CSG as <N, E>, where N is the extracted key API, which contains two types of nodes: server-side framework APIs and server-side key APIs; E represents the API dependency relationship, that is
[0018] Appendix Figure 2 is an example of CSG, which shows the process of constructing the semantic graph of the server-side key behavior of the present invention. Specifically, the framework API provided by the server (such as getLoginCode) is represented by an oval node, and the server-side key API (such as API1) is represented by a rectangular node; the edge represents the data dependency relationship between the nodes. Figure 2The process of obtaining and using keys is shown and includes corresponding semantic meanings. Specifically, in the entire key usage process, the applet first calls the server framework API (getLoginCode) to obtain the parameter loginCode; next, the server key API1 uses the parameter loginCode to obtain the encryption key encryptKey; thus, there is a data dependency relationship between the framework API (getLoginCode) and the server key API1. Similarly, after the framework API (getPhoneNum) obtains the encrypted phone number, the server key API2 uses the encryptKey obtained by API1 to decrypt the encrypted phone number. Therefore, there is a data dependency relationship between the server key API2, the framework API (getPhoneNum), and the server key API1.
[0019] The present invention constructs a set of semantic graphs according to the usage patterns of different keys, and these graphs will be used for further analysis.
[0020] (2) Client key behavior semantic understanding stage
[0021] Considering that keys do not have obvious characteristics in the applet client, to solve this problem, the present invention understands the client key semantics by performing data flow analysis on different types of applets. Specifically, it includes:
[0022] (1) Applet behavior analysis: Some super apps provide customized runtime engines (such as WeChat, Baidu, and Alipay), and the present invention can capture the code packages of applets. Therefore, the present invention performs data flow analysis to track network requests and extract potential important behaviors in the applet, such as obtaining loginCode from the server. However, some super apps run applets in the WebView environment. This type of applet needs to load code from its own server of the applet. To analyze these WebView-based applets, the present invention designs a dynamic crawler to obtain the client behaviors of the applets. In addition, the present invention also designs a new graph, called the client behavior graph, abbreviated as CBG, to represent the client semantics of the applet.
[0023] (2) Construction of the Client Behavior Graph: The Client Behavior Graph (CBG) is used to represent the behavior of the applet client. It contains two types of nodes: applet framework APIs and customized client APIs. The edges in the graph describe the data dependencies between these nodes. The key to constructing the CBG lies in establishing the data dependencies between different nodes, including explicit dependencies and implicit dependencies. Explicit dependencies refer to the dependencies formed through direct data propagation in the program code, usually manifested as variable assignments, function calls, parameter passing, etc. Such dependencies can be directly identified through data flow analysis. In contrast, implicit dependencies are not formed through direct data propagation but are indirectly reflected through global states, environment variables, etc. Such dependencies are usually difficult to directly identify through static analysis and require inference in combination with code semantics.
[0024] To extract explicit data dependencies, first perform data flow analysis on the applet client code. By deeply parsing key operations such as variable assignments, parameter passing, and function calls in the program code, trace the specific propagation path of data in the program, so as to accurately identify the explicit data dependencies within the program. Specifically, the present invention models specific APIs provided by super apps (i.e., platforms that integrate multiple functions and services and provide a runtime environment for a large number of applets, such as WeChat, Baidu, Alipay, etc.) and constructs an inter-procedural control flow graph for each file in the applet. Since many function calls in the applet are cross-file, inter-file analysis also needs to be performed. For inter-file function calls, the present invention connects the call edges between files by analyzing the "require" (module dependency) relationship between files to enhance the call graph. Then extract the parameter context information related to the key in the data flow to enrich the behavior semantics. Attached Figure 3 is an example of client semantics construction. Through inter-file analysis of the applet code, the dependencies in the data flow can be easily extracted. For example, if the return value of getLoginCode is further used by API 1’, there is an edge between these two APIs.
[0025] Since some super apps run mini programs in the WebView environment, it is impossible to directly obtain the code packages of such mini programs, and a dynamic method needs to be adopted to obtain the mini program code. Therefore, the present invention implements a dynamic analysis module to explore mini programs based on WebView. A common problem faced by traditional dynamic exploration is the coverage issue. They are difficult to summarize text input patterns and meet the unique constraints of text input. The present invention proposes a hint-guided dynamic exploration test for mini programs. Since mini program developers often provide hint information in the user interface to help users effectively fill in personal data, such as data type, data length, etc. Therefore, the present invention first obtains the GUI hierarchy in the current mini program interface and collects the elements near the input boxes to extract hint information. Then, predefined data is filled in the input boxes according to these hints. Finally, a value-based method is proposed to establish the correlation between different APIs. By comparing the parameter values and return values of different APIs, their dependency relationships are analyzed. At the same time, due to the diversity of mini program languages, the present invention is adapted to multiple languages, including English, Russian, Japanese, etc.
[0026] To extract implicit data dependencies, the present invention analyzes the development documents and models the patterns of implicit dependencies to supplement the CBG. For example, WeChat mini programs usually use wx.setStorageSync(key, value) to store necessary data and wx.getStorageSync(key) to obtain the stored data with the same key value, and these data cannot be directly tracked through the above steps. The observation of the present invention here is that most of these patterns appear in pairs, especially in the form of getter-setter patterns, such as wx.setstoragesync / wx.getStorageSync and wx.setStorage / wx.getStorage. Therefore, this problem can be transformed into identifying paired implicit dependencies between nodes. The present invention analyzes the development documents, identifies data-related operations and collects the usage patterns containing implicit data dependencies, and then the present invention uses a pattern analysis-based method to identify potential implicit dependencies on the same variable, thereby improving the CBG.
[0027] (III) Semantic-based similarity analysis stage
[0028] To detect whether there is a key leakage, it is necessary to determine whether the client behavior of the applet contains key semantic information, that is, whether the semantics of the CSG exists in the CBG. Since these semantics contain information at different levels, namely the behavior semantics of the server and the behavior semantics of the applet client. To bridge the semantic gap and discover whether there is key semantics in the applet client, the present invention turns to proving the semantic isomorphism between the CSG and the CBG, and designs an analysis algorithm based on semantic similarity to detect the key leakage problem in the applet.
[0029] Specifically, the present invention analyzes whether the CBG contains the semantic information of the CSG by identifying whether the CSG is a subgraph of the CBG. During the specific process of the applet using the key, the framework APIs provided by the super app remain unchanged, and the applet can directly call them on the client to obtain information, such as getLoginCode and getPhoneNum. Therefore, for such framework APIs, the corresponding API signature information can be directly matched. For other API operations, the present invention observes that operations with the same function have similar contexts or signatures (such as parameters and data dependencies) on the applet client and the server; in this regard, the present invention can match nodes according to the API signature information and associate these nodes based on similar signatures. For example, if the applet developer directly migrates the key API (involving the key in the parameters or return value) to the applet client, the present invention can easily identify the key leakage problem therein through API matching.
[0030] Based on the above observations, the present invention designs an algorithm that uses context information to bridge the semantic gap and prove the semantic isomorphism between the CSG and the CBG. This algorithm mainly consists of two steps - node reasoning and similarity analysis.
[0031] Node reasoning: First, it is necessary to find equivalent nodes in the CSG and the CBG. If these nodes are client operations, they can be directly matched according to the signature. To detect incorrect server operations, since the equivalent nodes have similar context or signature information, the present invention first analyzes the parameter information. For each parameter of the node, its source can be obtained according to the data dependency relationship (that is, the client behavior that generates the parameter). If the sources are the same (for example, the same framework API), it can be assumed that the parameters are also the same. Although most custom parameters have no clear source and even have confusion problems, the context information during the data propagation and use processes can assist in understanding the semantics of the parameters. For example, if a parameter comes from the custom API getPassword, its semantic information related to password can be extracted.
[0032] Therefore, the present invention first obtains the context of the parameters, including data dependencies (extracted according to data flow analysis) and parameter names. Then, the Levenshtein distance algorithm [1] is used to perform similarity analysis on the context information. Finally, if the nodes of CBG and CSG have the same source or similar context, the present invention associates these nodes.
[0033] Similarity analysis: The present invention compares the edges in the two graphs of CBG and CSG (i.e., the data dependencies between nodes) to perform similarity analysis on the semantics of the mini-program key behavior. Specifically, a depth-first traversal strategy is adopted to explore the edges in CSG, and it is checked whether there are dependencies between the corresponding nodes in CBG. In each iteration, an edge is selected from CSG, and the corresponding nodes in CBG are traversed (obtained in the node inference stage). Then, it is checked whether there is an edge between these corresponding nodes.
[0034] Since there may be multiple key leakage paths in the mini-program, there may be multiple corresponding nodes in CBG. As long as there is a data dependency between any pair of nodes, it is considered that there is an edge. When the nodes are not directly dependent (there are other nodes), they can be associated according to the data dependency. If all the edges in CSG match the edges in CBG, the similarity analysis is completed, and it can be inferred that the mini-program has leaked the key corresponding to CSG. Since multiple keys may be leaked in the mini-program, the present invention will match all semantic graphs to check whether there are other leakage situations.
[0035] The beneficial effects of the present invention are:
[0036] The present invention proposes a new key leakage detection method for static information flow analysis technology for the semantics of mini-program key behavior, which effectively detects key leakage problems in mini-programs based on the similarity of key semantics between the client and the server. Thus, it provides more comprehensive monitoring and protection for mini-program key management, and effectively reduces the leakage risk of key protected critical information and sensitive resources. Brief Description of the Drawings
[0037] Figure 1 It is the overall architecture diagram of the detection system.
[0038] Figure 2 It is an example of the key usage semantic graph. Among them, a is the mini-program development document, and b is CSG.
[0039] Figure 3 It is an example of client semantic construction. Among them, a is the mini-program client code, and b is CBG.
[0040] Figure 4 It is an example of the mini-program data dependency code.
[0041] Figure 5 This is an example of semantic similarity analysis. Here, a is CSG and b is CBG. Specific implementation manner
[0042] The present invention will be further explained below through specific examples in conjunction with the accompanying drawings.
[0043] (1) Server-side key behavior semantic understanding stage
[0044] The present invention has written a Python script to implement the functions of collecting and analyzing the development documents of mini-programs in the above design for 6 well-known domestic and foreign mini-program platforms. In order to quickly extract structured document information, the present invention uses XPath to construct the structural patterns of the development documents of different mini-program platforms, and then automatically parses the API semantic information therein. Figure 2 (a) shows an example of a mini-program development document, where the API information is structured and organized in different modules. For example, the URL address of the API is located in Request Address, and the parameter information is located in RequestParameters. Therefore, XPath can be used to locate specific elements and extract API information.
[0045] At the same time, the present invention further analyzes the dependency relationships between different APIs. As shown in Figure 2 (a), it can be observed that the description information of the parameter code of API 1 contains a hyperlink pointing to API getLoginCode. Therefore, the present invention can construct the dependency relationship between these two APIs. In addition, by analyzing the similarity of variable naming between different APIs, it can be found that the return value of API1 and the parameter of API 2 both contain a variable named encryptKey. Therefore, it can be considered that there is a dependency relationship between API 1 and API 2.
[0046] Based on the extracted API information and the dependency relationships between different APIs, the present invention can construct a semantic graph of key usage as shown in Figure 2 (b). Finally, the present invention has collected the document information of 1085 server-side APIs in total. And according to the key types of different mini-program platforms, the APIs related to keys are screened out, and then 658 semantic graphs are constructed according to the usage patterns of different keys for further analysis in the future.
[0047] (2) Client-side key behavior semantic understanding stage
[0048] In the stage of mini-program behavior analysis, the present invention first obtains code packages of different types of mini-programs. For mini-programs based on WebView, the present invention designs a dynamic crawler, explores the mini-program by simulating user interactions based on Android UI Automator, dynamically obtains the client behavior of the mini-program, and then crawls the mini-program code package. During the dynamic exploration process, the present invention fills in according to the prompt information near the input box. For example, if the prompt information of the input box is "name", the present invention will automatically fill in the preset name information.
[0049] For the extracted mini-program code package, the static analysis tool JAW is used to perform program analysis on the mini-program code. JAW is a graphical security analysis framework for client-side JavaScript. It can design and execute customized security-related program analysis, including data flow analysis between predefined JavaScript sources and receivers, control flow and reachability analysis, pattern matching through the abstract syntax tree (AST), etc. Therefore, it is selected for the static analysis of the mini-program.
[0050] Appendix Figure 4 shows the process of extracting the dependency relationship in the mini-program code. The present invention first extracts the explicit data dependency relationship in the mini-program code based on data flow analysis. For example, we can observe that the data t on line 9 is propagated to line 11, and the data sk on line 16 is propagated to line 22. Therefore, these explicit dependency relationships can be extracted by tracing the data flow. Then, the APIs related to implicit dependencies are analyzed, and the corresponding API list (see Table 1) is sorted out. Furthermore, based on these APIs, the implicit data dependency relationship in the mini-program is analyzed.
[0051] Table 1
[0052] API related to implicit dependencies Code example setStorageSync setStorageSync("name","Ja**") getStorageSync value = getStorageSync("name") setStorage setStorage({key: "name", data: "Ja**"}) getStorage getStorage({key: "name", success(res){…}}) setData setData({"name": "Ja**"}) globalData getApp().globalData.name = "Ja**" 。
[0053] As shown in the appendix Figure 4 As shown, in line 11, t.data.sk is stored in "sk" through the wx.setStorageSync API, and then obtained through the wx.getStorageSync API in line 16. Although there is no explicit data propagation relationship, through the positioning and analysis of the privacy dependency-related APIs we collected, the present invention can extract such implicit data dependency relationships.
[0054] Finally, based on the analysis of the mini-program code behavior, the present invention can construct a semantic graph of the mini-program client key behavior as shown in Appendix Figure 3 (b).
[0055] (III) Semantic-based similarity analysis
[0056] The present invention has written a Python script to implement node inference and similarity analysis functions in the process of semantic-based similarity analysis.
[0057] In node inference, the present invention uses the Levenshtein distance algorithm for similarity analysis. This algorithm can quickly calculate the minimum edit distance between strings and is used in many works. Therefore, it is selected to implement the similarity analysis module in the present invention. For the parameter setting of the similarity analysis algorithm, in order to improve the accuracy of node inference, preliminary research has been carried out to select the best similarity threshold for evaluation. The similarity threshold is set to 0.9 to compare the context similarity while maintaining a low false alarm rate. The detailed process of node inference is as shown in the appendix Figure 5 shown. Among them, API 1 in CSG and API 1' in CBG accept the same parameters from getLoginCode and have similar context information. Therefore, we can associate these two nodes.
[0058] In similarity analysis, the present invention determines its matching by comparing the edges (i.e., the data dependency relationships between nodes) in two graphs. As shown in the appendix Figure 5 shown, for the similar node pairs obtained in the node inference stage, the present invention determines whether there is an edge by analyzing the data dependency relationships between them. For example, since there is a data dependency relationship between API 1 and API 2, there is an edge between them; similarly, there is a data dependency relationship between API 1' and API 2', so there is also an edge. After constructing all the edges, it can be found that each edge in CSG can match the edge in CBG, which indicates that CSG is a subgraph of CBG. Therefore, it can be inferred that the key corresponding to CSG has been leaked.
[0059] References
[0060] [1] Karin Beijering, Charlotte Gooskens, Wilbert Heeringa. Predicting intelligibility and perceived linguistic distance by means of the levenshtein algorithm[J]. Linguistics in the Netherlands, 2008, 25(1), 13–24. DOI: 10.1075 / avt.25.05bei.
Claims
1. A method for detecting dynamic unstructured key leakage in a mini program, characterized in that: The semantic analysis of mini-program key behavior is combined with traditional static information flow analysis technology to detect the leakage of dynamic unstructured keys in mini-programs. Here, the key types are summarized into three categories: (1) Root key, which is used to generate and manage other keys; (2) Access key, which is used to verify the user's resource access rights; (3) Encryption key, a key used for data encryption; For different types of keys, their different behavioral semantics are analyzed and detected. The detection process is divided into three stages: (I) The server-side key behavior semantic understanding stage includes: (1) Mini Program Development Document Collection: First, the mini program development documents are automatically parsed to obtain the semantic information contained in each server-side application programming interface (API), including the API name, parameters, and function description. Parsing each developer document can summarize and organize a fixed structural pattern, which can then be expanded to the automated analysis of other APIs and large-scale crawling of document information. (2) Analysis of mini program development documents: Based on the semantic information contained in the API crawled above, the key usage semantics of the server is understood; in order to facilitate mini program developers to implement server operations, the mini program development document describes how to obtain the parameters in the API description; specifically, a hyperlink pointing to the source of the parameter is directly attached to the parameter description of an API; Therefore, a heuristic strategy is designed to extract the data dependency between server-side APIs. Through this strategy, different API operations are linked to each other to build a complete API behavior sequence. Specifically, starting from the entry API, the relevant dependent server-side APIs are identified according to their parameter descriptions. This process is iterated until no new dependent APIs are found. After the above steps, considering that the description information of the development document may lack corresponding hyperlinks, in order to fully restore the data dependency between server-side APIs, the naming similarity of variables is further used to establish the data dependency between different APIs. Based on this similar naming, the data dependency between different APIs is analyzed. (3) Construction of key usage semantic graph: In order to represent API-level semantics and take into account the correlation between server-side APIs, a key usage semantic graph, abbreviated as CSG, is designed. The key usage semantic graph is used to represent the semantics of the mini program server, which includes key-related operations. CSG is defined as<N,E> , where N is the extracted key API, which contains two types of nodes: server framework API and server key API; E represents the API dependency, i.e. Build a collection of semantic graphs based on the usage patterns of different keys, which will be used for further analysis; (II) Client key behavior semantic understanding stage Considering that the key has no obvious characteristics in the mini program client, we analyze the data flow of different types of mini programs to understand the client key semantics. Specifically, we: (1) Mini-program behavior analysis: Provide a customized runtime engine for super apps to capture the code packages of mini-programs; perform data flow analysis to track network requests and extract potentially important behaviors in mini-programs; for some mini-programs running in the WebView environment of super apps, load the code from the mini-program’s own server; in order to analyze these WebView-based mini-programs, design a dynamic crawler to obtain the client behavior of the mini-programs; (2) Construction of client behavior graph: The client behavior graph, abbreviated as CBG, is used to represent the behavior of the mini program client. It contains two types of nodes: the mini program framework API and the customized client API. The edges in the graph describe the data dependencies between these nodes. The construction of CBG is to establish data dependencies between different nodes, including explicit dependencies and implicit dependencies. Explicit dependencies refer to dependencies formed through direct data propagation in the program code, which are manifested as variable assignments, function calls, and parameter passing. This type of dependency is directly identified through data flow analysis. In contrast, implicit dependencies are not formed through direct data propagation, but are indirectly reflected through global states and environment variables. This type of dependency is difficult to directly identify through static analysis and needs to be inferred in combination with code semantics.
3. Semantic-based similarity analysis In order to detect whether there is a key leakage, it is necessary to determine whether the client behavior of the mini program contains key semantic information, that is, whether the semantics of CSG exist in CBG; since these semantics contain information at different levels, namely the behavioral semantics of the server and the behavioral semantics of the mini program client; in order to bridge the semantic gap and find out whether there is key semantics in the mini program client, we turn to proving the semantic isomorphism between CSG and CBG, and design an analysis algorithm based on semantic similarity to detect key leakage problems in mini programs; Specifically, by identifying whether CSG is a subgraph of CBG, we analyze whether CBG contains the semantic information of CSG; in the process of the mini program using the key, the framework API provided by the super application remains unchanged, and the mini program can directly call it on the client to obtain information; for such framework APIs, the corresponding API signature information is directly matched; for other API operations, it is observed that operations with the same functions have similar contexts or signatures on the mini program client and server; in this regard, nodes are matched according to the API signature information, and these nodes are associated based on similar signatures; if the mini program developer directly migrates the key API (involving keys in parameters or return values) to the mini program client, the key leakage problem can be easily identified through API matching.
2. The method for detecting dynamic unstructured key leakage in a mini-program according to claim 1, characterized in that: During the client key behavior semantic understanding phase: In order to extract explicit data dependencies, we first perform data flow analysis on the mini program client code. By deeply analyzing the key operations of variable assignment, parameter passing, and function call in the program code, we track the specific propagation path of data in the program, thereby accurately identifying the explicit data dependencies within the program. In order to extract implicit data dependencies, development documents are analyzed and the patterns of implicit dependencies are modeled to supplement CBG. By analyzing development documents, data-related operations are identified and usage patterns containing implicit data dependencies are collected. Then, a pattern analysis-based method is used to identify potential implicit dependencies on the same variable, thereby improving CBG.
3. The method for detecting dynamic unstructured key leakage in a mini-program according to claim 1, characterized in that: In the semantic-based similarity analysis phase, an algorithm is designed to use contextual information to bridge the semantic gap and prove the semantic isomorphism between CSG and CBG; the algorithm is divided into node reasoning and similarity analysis; Node reasoning: First, we need to find equivalent nodes in CSG and CBG; If these nodes are client operations, they are matched directly based on the signature. First, the parameter information is analyzed. For each parameter of the node, its source is obtained based on the data dependency, that is, the client behavior that generates the parameter. If the sources are the same, the parameters are assumed to be the same; Specifically, we first obtain the context of the parameters, which includes data dependencies and parameter names. Then we use the Levenshtein distance algorithm to perform similarity analysis on the context information. Finally, if the nodes of CBG and CSG have the same source or similar context, we associate these nodes. Similarity analysis: Compare the edges in the CBG and CSG graphs to perform similarity analysis on the semantics of the mini-program key behavior. Specifically, a depth-first traversal strategy is used to explore the edges in the CSG and check whether there is a dependency relationship between the corresponding nodes in the CBG. In each iteration, an edge is selected from CSG and the corresponding nodes in CBG are traversed; then, it is checked whether there is an edge between these corresponding nodes; Since there may be multiple key leakage paths in the mini program, there may be multiple corresponding nodes in CBG. As long as there is a data dependency relationship between any pair of nodes, an edge is considered to exist; when there is no direct dependency between the nodes, they can be associated based on the data dependency relationship; if all edges in CSG match the edges in CBG, the similarity analysis is completed and it is inferred that the mini program leaked the key corresponding to CSG; since multiple keys may be leaked in the mini program, all semantic graphs will be matched to check whether there are other leakage situations.