A Small Program Cloning Detection Method Based on Complex Network Analysis

Through the method based on complex network analysis, the statistical and layout features of the applet are extracted and the mini program cloning detection is combined with machine learning, which solves the problem of insufficient speed and accuracy in the existing technology, and efficient and accurate cloning detection is achieved, improving user security.

CN116028112BActive Publication Date: 2025-07-25XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310045745.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-30
Publication Date
2025-07-25
Estimated Expiration
2043-01-30

AI Technical Summary

Technical Problem

The existing mini-program cloning detection methods cannot take into account the detection speed and detection accuracy, resulting in high false alarm rates and missed alarm rates, which cannot effectively ensure user safety.

Method used

Using a method based on complex network analysis, the coarse-grained statistical features and fine-grained layout features and code features are extracted from the source code through static analysis, and a classifier is used to determine whether the applet is cloned, including preprocessing, extraction of statistical features, layout features, file-dependent features and dual-layer dependency features, and binary classification is performed in combination with machine learning.

Benefits of technology

It improves the speed of mini-program cloning detection, reduces the false alarm rate and missed alarm rate, and improves the security of users when using mobile applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116028112B_ABST
    Figure CN116028112B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting applet cloning based on complex network analysis to solve the current problem of applet cloning detection. This method first preprocesses the applet to be detected, and extracts coarse-grained statistical features, as well as fine-grained layout features and code features from the source code through static analysis. According to its statistical features, it is divided into different cloned applet family clusters. According to the layout features and code features, the similarity vector between the applet to be detected and each applet in the cloned applet family cluster is calculated, and a classifier is used to determine whether two applets are cloned. Through the above method, the cloning situation of applets can be detected, the false positive rate and false negative rate are reduced, and the security guarantee for users when using mobile applications is improved. A new method for applet cloning detection is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of program clone detection in mobile applications, and particularly to a method for detecting mini-program clones based on complex network analysis. Background Art

[0002] Plagiarism will seriously infringe the rights of original developers. For example, replacing advertising links to seek economic benefits, or confusing the public to divert traffic to malicious mini-programs. Such behavior will harm the mini-program ecosystem. At the same time, plagiarists may also introduce malicious code to carry out malicious behaviors such as stealing user privacy.

[0003] Currently, code clone detection techniques can be mainly divided into four categories: text-based, lexical-based, syntax-based, and semantic-based.

[0004] Text-based detection methods are divided into two types. One is to regard it as a string similarity problem and only compare from characters; the other is to extract fine-grained features from the character perspective, such as special characters. Text-based detection is fast, but the accuracy is low and it is difficult to resist obfuscation.

[0005] Lexical-based detection methods can also be divided into two methods: Token-based detection methods and API-based detection methods. The Token-based detection method processes identifiers in the code and returns Tokens for source code comparison, which can effectively solve the problem of variable name and function name replacement. Since APIs are functions provided by the system or framework to developers in advance, no matter how the plagiarist obfuscates the control flow and data flow of the code, the most basic API calls will not change significantly. The API-based detection method usually extracts the number of API calls as a feature. The accuracy of lexical-based detection is slightly higher than that of text-based detection, but due to the lack of syntax and semantic analysis, it is prone to misjudgment.

[0006] Syntax-based detection usually believes that similar code segments also have similar syntax structures. This method will use a parser to parse the target source code into an AST abstract syntax tree, and then compare the tree structures between codes or extract features from the tree structures to compare the similarity of the trees. The source code segments corresponding to the similar subtrees are the cloned codes. Syntax-based detection has a large computational overhead, but high accuracy.

[0007] Semantic-based detection methods will analyze the control flow graph and data flow graph from the source code, and then construct a PDG in combination with the control flow graph and data flow graph. There are also those that only rely on the control flow graph and data flow graph for detection. Use the subgraph isomorphism algorithm or the algorithm for extracting graph features to compare graph similarity. The source code corresponding to the similar graphs is the cloned code. Semantic-based detection has strong anti-obfuscation ability and high accuracy, but also has a large computational overhead. Summary of the Invention

[0008] The content of the present invention is to propose a mini-program cloning detection method based on complex network analysis to solve the problem that the single method of mini-program cloning detection cannot take into account both detection speed and detection accuracy. This method first preprocesses the mini-program to be detected, and extracts coarse-grained statistical features SF, fine-grained layout features LF, and code-dimensional features, namely custom function features CFF, file dependency features FDF, and double-layer dependency features TLDF from the source code through static analysis. According to its statistical features SF, it is divided into different cloned mini-program family clusters. According to the layout features LF, custom function features CFF, file dependency features FDF, and double-layer dependency features TLDF, the similarity vectors between the mini-program to be detected and each mini-program in the cloned mini-program family cluster are calculated, and the classifier judges whether two mini-programs are cloned. Through the above method, the cloning situation of mini-programs can be detected, the detection speed is accelerated, the false positive rate and false negative rate are reduced, and the security guarantee for users when using mobile applications is improved. It makes up for the blank of the mini-program cloning detection method.

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] Step S101: Preprocess the mini-program S to be detected, including decompilation, de-obfuscation, and extraction of the main package;

[0011] Step S102: According to the source code of the preprocessed mini-program S to be detected obtained in step S101, extract statistical features SF, layout features LF, custom function features CFF, file dependency features FDF, and double-layer dependency features TLDF by analyzing the file type and the source code abstract syntax tree;

[0012] Step S103: Calculate the distance between the statistical features SF of the mini-program S to be detected obtained in step S102 and the center of the cloned mini-program family cluster, and divide the mini-program S to be detected into the nearest cloned mini-program family cluster;

[0013] Step S104: According to the cloned mini-program family cluster obtained in step S103, form mini-program pairs by combining the mini-program S to be detected with each mini-program in the cloned mini-program family cluster. The similarity vectors of the mini-program pairs are the similarities of the layout features LF, custom function features CFF, file dependency features FDF, and double-layer dependency features TLDF respectively;

[0014] Step S105: Use the similarity vectors of the mini-program pairs with pre-labeled tags of cloned and non-cloned, and use machine learning methods to construct a classifier;

[0015] Step S106: According to the classifier obtained in step S105, input the similarity vector of the applet pair composed of the applet S to be detected in step S104 and each applet in the cloned applet family cluster, and perform binary classification on the applet pair to be divided into cloned applet pairs and non-cloned applet pairs, thereby finding the applet cloned from the applet S to be detected.

[0016] Further, the specific step S101 is as follows:

[0017] Step S201: Decompile the packaged file of the applet S to be detected according to the applet packaging rule to obtain its source code;

[0018] Step S202: According to the source code of the applet S to be detected obtained in step S201, extract the obfuscation pattern and match it with the existing obfuscation methods. If it can match the existing obfuscation methods, use the corresponding de-obfuscation method to de-obfuscate the source code;

[0019] Step S203: According to the de-obfuscated source code of the applet S to be detected obtained in step S202, collect and organize to obtain the commonly used third-party library T of the applet. For the file f in the source code S of the applet to be detected, use the file name and file size attributes to match with the known third-party library files in T. If the file f in the source code S meets the match, filter out the third-party library file for f.

[0020] Further, the statistical feature SF in the step S102 is a 241-dimensional vector. The first 3 dimensions are the number of pages of the applet, the average number of static files, and the number of developer-defined functions respectively. The 4th dimension to the 137th dimension are the source API call times, and the 138th dimension to the 241st dimension are the sink API call times.

[0021] Further, the specific extraction of the statistical feature SF in the step S102 is as follows:

[0022] Step S301: According to the preprocessed source code of the applet S to be detected obtained in step S101, count the number of pages of the applet from the routing components, routing functions, pages field and tabbar field of app.json;

[0023] Step S302: According to the preprocessed source code of the applet S to be detected obtained in step S101, count the average number of static resource files in the static resource file directory by comparing the file types. The static resource file directory refers to the directory where all files in the directory are static resource files, and the static resource file refers to the file that is not json, js, class html, and css;

[0024] Step S303: According to the preprocessed source code of the to-be-detected mini-program S obtained in Step S101, count the number of functions defined by the developer. The functions defined by the developer refer to functions that are not system-defined or third-party library-defined.

[0025] Step S304: According to the preprocessed source code of the to-be-detected mini-program S obtained in Step S101, count the number of source API calls and sink API calls. Source API is an API that returns sensitive data as a return value, and sink API is an API that passes sensitive data as a parameter.

[0026] Step S305: Combine the data obtained in Step S301, Step S302, Step S303, and Step S304 to form a feature vector as the statistical feature SF of the to-be-detected mini-program S.

[0027] Further, the layout feature LF in Step S102 is a hash sequence, denoted as LF(S) = <fh1,…,fh N >, fh i represents the i-th hash value.

[0028] Further, the specific process of extracting the layout feature LF in Step S102 is as follows:

[0029] Step S401: According to the preprocessed source code of the to-be-detected mini-program S obtained in Step S101, parse the class html file for page layout display, and extract the component sequence.

[0030] Step S402: According to the component sequence obtained in Step S401, use weak hashing to slice it, such as Alder32. The slicing condition is that the remainder of the weak hash value within the slice modulo 64 is 63.

[0031] Step S403: According to the sliced component sequence obtained in Step S402, use strong hashing to calculate the hash value of the slice, such as FNV64, and use the new hash sequence as its layout feature LF.

[0032] Further, the custom function feature CFF in Step S102 is a hash sequence, denoted as CFF(S) = <fh1,…,fh N >, fh i represents the i-th hash value.

[0033] Further, the specific process of extracting the custom function feature CFF in Step S102 is as follows:

[0034] Step S501: According to the preprocessed source code of the to-be-detected mini-program S obtained in Step S101, parse the js file for executing the functional logic into an AST abstract syntax tree.

[0035] Step S502: According to the AST abstract syntax tree of the js file obtained in Step S501, extract the type sequence of the developer-defined functions. The type sequence refers to the sequence obtained by pre-order traversing the AST abstract syntax tree and sequentially taking out the type attribute values of the nodes.

[0036] Step S503: According to the component sequence obtained in Step S502, use weak hashing to slice it. The slicing condition is that the remainder of the weak hash value within a slice modulo 64 is 63.

[0037] Step S504: According to the sliced component sequence obtained in Step S503, use strong hashing to calculate the hash value of the slice, and use the new hash sequence as its custom function feature CFF.

[0038] Furthermore, the file dependency feature FDF in Step S102 is a file dependency graph, denoted as FDF(S) = (FeileV(S), FeileE(S)). FeileV(S) is the set of nodes of the file dependency graph of the mini-program S to be detected, and the node attribute is the file name. FeileE(S) is the set of edges of the file dependency graph of the mini-program S to be detected, and FeileE(S) = {<v i ,v k >} where v i ,v k ∈FeileV(S).

[0039] Furthermore, the specific process of extracting the file dependency feature FDF in Step S102 is as follows:

[0040] Step S601: Take all js files under the source code file directory of the mini-program S to be detected as the set of nodes of the file dependency graph.

[0041] Step S602: Traverse the js files, parse them into AST abstract syntax trees, extract the import dependency statements from the AST abstract syntax trees, and if the imported file is in the set of nodes of the file dependency graph, add this pair of files to the set of edges of the file dependency graph.

[0042] Furthermore, the two-layer dependency feature TFDF in Step S102 is a two-layer graph structure. The upper layer graph is a file dependency graph, and the lower layer graph is a function call graph, denoted as FDF(S) = (FeileV′(S), FileE′(S)). FileV′(S) is the set of nodes of the file dependency graph of the mini-program S to be detected, and the node attribute is the function call graph of this file. FileE′(S) is the set of edges of the file dependency graph of the mini-program S to be detected, and FileE′(S) = {<v i ,v k >} where v i ,v k ∈FeileV′(S).

[0043] Further, the specific process of extracting the two-layer dependency feature TFDF in step S102 is as follows:

[0044] Step S701: Use all js files in the source code file directory of the to-be-detected applet S as the node set of the file dependency graph;

[0045] Step S702: Traverse the js files, parse them into AST abstract syntax trees, extract the function call relationships within the files from the AST abstract syntax trees to construct function call graphs, and use the function call graphs as the attributes of the js files;

[0046] Step S703: According to the AST abstract syntax tree obtained in step S702, extract the import dependency statements. If the imported file is in the node set of the file dependency graph, add this pair of files to the edge set of the file dependency graph.

[0047] Further, in step S103, the center of the cloned applet family cluster is pre-solved, that is, the average value of the statistical features SF of the applets within the cloned applet family cluster is calculated, the distance between the statistical feature SF of the to-be-detected applet S and the center of each cloned applet family cluster is calculated, and it is classified into the nearest cloned applet family cluster.

[0048] Further, in step S104, the similarity calculation method of the layout feature LF of the applet pair is to traverse the layout features LF of both, regard the hash value as a string, calculate the Levenshtein ratio to compare the similarity of the two hash values. If the Levenshtein ratio is greater than a certain threshold, the number of similar hash values is incremented by one, and the number of similar hash values is divided by the maximum value of the lengths of the layout features LF of the two applets; the custom function feature CFF of the applet is also a hash sequence, and the same similarity calculation method as the layout feature LF is used.

[0049] Further, in step S104, the similarity calculation method of the file dependency feature FDF of the applet pair is to use the Weisfeiler-Lehman algorithm to iterate the file dependency graphs of both, and calculate the similarity according to the number of similar labels in the output label set.

[0050] Further, in step S104, since the node attribute of the two-layer dependency feature TLDF of the applet pair is a function call graph, the similarity calculation method is to first traverse the file dependency graphs of both, use the Weisfeiler-Lehman algorithm to calculate the similarity of the function call graphs of the nodes. If it is greater than a certain threshold, the node is put into the anchor point set, and the number of nodes in the anchor point set is divided by the maximum value of the number of nodes in the file dependency graphs of both.

[0051] Further, the similarity vector of the mini-program pair in step S105 is the similarity of layout features LF, custom function feature CFF, file dependency feature FDF, and two-layer dependency feature TLDF of two mini-programs. If two mini-programs are similar in function and layout after manual review and there are official forums or news reports alleging plagiarism between them, they are considered to have a cloning relationship and are labeled as a cloned mini-program pair; otherwise, they are a non-cloned mini-program pair. The input of the classifier is the similarity vector of the mini-program pair, and the output is 0 and 1. 0 represents that the mini-program pair is a non-cloned mini-program pair, and 1 represents that the mini-program pair is a cloned mini-program pair.

[0052] A further improvement of the present invention is that: the layout feature extraction method in step S102 is to parse the wxml file to extract the component sequence, and then convert the component sequence into a fuzzy hash sequence, that is, first use weak hashing to shard and then use strong hashing to calculate the shard hash value, and splice the shard hash values.

[0053] A further improvement of the present invention is that: in the two-layer dependency feature in step S102, the function call graph is used as an attribute of the file dependency graph node. The nodes of the function call graph are the functions of the file, and the edges are the function call relationships of the file.

[0054] A further improvement of the present invention is that: in step S103, the to-be-detected mini-program is classified into different cloned mini-program family clusters according to its statistical features, and in step S106, it is judged whether the two are cloned by using a classifier based on the similarity vector of the mini-program pair.

[0055] Compared with the prior art, the present invention has the following advantages:

[0056] 1) The present invention uses coarse-grained features for clustering first and then calculates the similarity vector of fine-grained features for analysis, which improves the efficiency compared with directly calculating the similarity vector;

[0057] 2) The layout features proposed by the present invention use the combination of weak hashing and strong hashing, and the code features use the file dependency graph and the function call graph, which can effectively resist code obfuscation and improve the robustness and accuracy;

[0058] 3) The present invention uses a classifier in machine learning to judge mini-program cloning, which is more accurate than manually setting thresholds and has better generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is the overall flowchart of the mini-program cloning detection method based on complex network analysis of the present invention;

[0060] Figure 2 is the method flowchart for extracting layout features from the source code of the to-be-detected mini-program through static analysis of the present invention;

[0061] Figure 3 This is the flowchart of the method for extracting custom function features from the source code of the applet to be detected in the present invention through static analysis. Specific embodiments

[0062] The following will detail the specific embodiments of the applet cloning detection method based on complex network analysis of the present invention with reference to the accompanying drawings.

[0063] Figure 1 This is the overall flowchart of the applet cloning detection method based on complex network analysis of the present invention;

[0064] The present invention discloses an applet cloning detection method based on complex network analysis, including the following steps:

[0065] Step S101: Preprocess the applet S to be detected, including decompilation, deobfuscation, and extraction of the main package.

[0066] Specifically, it can be divided into the following steps:

[0067] Step S201: Decompile the packaged file of the applet S to be detected according to the applet packaging rule to obtain its source code;

[0068] Step S202: Extract the obfuscation pattern from the source code of the applet S to be detected obtained in step S201 and match it with the existing obfuscation methods. If it can be matched with the existing obfuscation methods, use the corresponding deobfuscation method to deobfuscate the source code;

[0069] Step S203: According to the deobfuscated source code of the applet S to be detected obtained in step S202, collect and organize to obtain the commonly used third-party libraries T of the applet. For the file f in the source code S of the applet to be detected, match it with the known third-party library files in T using the file name and file size attributes. If the file f in the S source code meets the match, filter out the third-party library file for f.

[0070] Step S102: According to the preprocessed source code of the applet S to be detected obtained in step S101, extract statistical features SF, layout features LF, custom function features CFF, file dependency features FDF, and two-layer dependency features TLDF by analyzing the file type and the source code abstract syntax tree;

[0071] Specifically, the statistical feature SF is a 241-dimensional vector. The first 3 dimensions are respectively the number of pages of the applet, the average number of static files, and the number of developer-defined functions. The 4th dimension to the 137th dimension are the source API call times, and the 138th dimension to the 241st dimension are the sink API call times. The extraction of statistical features can be divided into the following steps:

[0072] Step S301: Based on the preprocessed source code of the mini program S to be detected obtained in step S101, count the number of pages of the mini program from the routing components, routing functions, pages field and tabbar field of app.json. The routing component is a component with page jump function such as <navigator>, the routing function is a function in the API that has the page jump function, such as navigateTo;

[0073] Step S302: According to the preprocessed source code of the to-be-detected applet S obtained in step S101, calculate the average number of static resource files in the static resource file directory by comparing the file types. The static resource file directory refers to a directory where all files are static resource files, and static resource files refer to files that are not json, js, class html, and css files;

[0074] Step S303: According to the preprocessed source code of the to-be-detected applet S obtained in step S101, count the number of functions defined by the developer himself. Functions defined by the developer himself refer to functions that are not system-defined and third-party library-defined functions;

[0075] Step S304: According to the preprocessed source code of the to-be-detected applet S obtained in step S101, count the number of source API calls and sink API calls. Source API is an API that returns sensitive data as a return value, and sink API is an API that passes sensitive data as a parameter;

[0076] Step S305: Combine the data obtained in steps S301, S302, S303, and S304 into a feature vector as the statistical feature SF of the to-be-detected applet S.

[0077] Step S103: According to the preprocessed source code of the to-be-detected applet S obtained in step S101, use static analysis to extract the layout feature LF;

[0078] The layout feature LF is a hash sequence, expressed as LF(S) = <fh1,…,fh N >, fh i represents the i-th hash value.

[0079] Figure 2 This is the flowchart of the method for extracting layout features from the source code of the to-be-detected applet by static analysis in the present invention.

[0080] Specifically, it can be divided into the following steps:

[0081] Step S401: According to the preprocessed source code of the to-be-detected applet S obtained in step S101, parse the class html file for page layout display, and extract the component sequence;

[0082] Step S402: According to the component sequence obtained in step S401, use weak hashing to slice it. Weak hashing uses Alder32, and the slicing condition is that the remainder of the weak hash value within the slice divided by 64 is 63;

[0083] Step S403: According to the sharded component sequence obtained in step S402, use strong hashing to calculate the hash value of the shard, and use the new hash sequence as its layout feature LF. The strong hashing uses FNV64.

[0084] The custom function feature CFF is a hash sequence, expressed as CFF(S) = <fh1, …, fh N >],fh i represents the i-th hash value.

[0085] Figure 3 This is the flowchart of the method for extracting custom function features from the source code of the to-be-detected applet through static analysis in the present invention.

[0086] Specifically, it can be divided into the following steps:

[0087] Step S501: According to the preprocessed source code S of the to-be-detected applet obtained in step S101, parse the js file for executing the function logic into an AST abstract syntax tree;

[0088] Step S502: According to the AST abstract syntax tree of the js file obtained in step S501, extract the type sequence of the developer-defined functions. The type sequence refers to the sequence formed by taking out the type attribute values of the nodes in turn by pre-order traversing the AST abstract syntax tree;

[0089] Step S503: According to the component sequence obtained in step S502, shard it using weak hashing. The weak hashing uses Alder32, and the sharding condition is that the remainder of the weak hash value within the shard divided by 64 is 63;

[0090] Step S504: According to the sharded component sequence obtained in step S503, use strong hashing to calculate the hash value of the shard. The strong hashing uses FNV64, and use the new hash sequence as its custom function feature CFF.

[0091] The file dependency feature FDF is a file dependency graph, expressed as FDF(S) = (FileV(S), FileE(S)), where FileV(S) is the set of nodes of the file dependency graph of the to-be-detected applet S, and the node attribute is the file name. FileE(S) is the set of edges of the file dependency graph of the to-be-detected applet S, and FileE(S) = {<v i , v k >} where v i , v k ∈FileV(S). Specifically, the extraction steps can be divided into:

[0092] Step S601: Take all js files under the source code file directory of the to-be-detected applet S as the set of nodes of the file dependency graph;

[0093] Step S602: Traverse the js files, parse them into an AST (Abstract Syntax Tree), extract the import dependency statements from the AST. If the imported file is in the file dependency graph node set, add this pair of files to the file dependency graph edge set.

[0094] The two-layer dependency feature TFDF is a two-layer graph structure. The upper layer graph is the file dependency graph, and the lower layer graph is the function call graph, denoted as FDF(S) = (FileV′(S), FileE′(S)). FileV′(S) is the file dependency graph node set of the mini-program S to be detected, and the node attribute is the function call graph of this file. FileE′(S) is the file dependency graph edge set of the mini-program S to be detected, and FileE′(S) = {<v i ,v k >} where v i ,v k ∈FileV′(S). Specifically, the extraction steps can be divided into:

[0095] Step S701: Take all js files under the source code file directory of the mini-program S to be detected as the file dependency graph node set;

[0096] Step S702: Traverse the js files, parse them into an AST, extract the function call relationships within the files to construct a function call graph, and take the function call graph as the attribute of this js file;

[0097] Step S703: According to the AST obtained in Step S702, extract the import dependency statements. If the imported file is in the file dependency graph node set, add this pair of files to the file dependency graph edge set.

[0098] Step S103: According to the statistical feature SF of the mini-program S to be detected obtained in Step S102, calculate its distance from the center of the clone mini-program family cluster, and divide the mini-program S to be detected into the nearest clone mini-program family cluster;

[0099] Specifically, in Step S105, the center of the clone mini-program family cluster is pre-solved, that is, the average value of the statistical features SF of the mini-programs within the clone mini-program family cluster is calculated, the distance between the statistical feature SF of the mini-program S to be detected and the center of each clone mini-program family cluster is calculated, and it is divided into the nearest clone mini-program family cluster.

[0100] Step S104: According to the clone mini-program family cluster obtained in Step S105, form mini-program pairs by pairing the mini-program S to be detected with each mini-program in the clone mini-program family cluster, calculate the similarity of the layout feature LF, custom function feature CFF, file dependency feature FDF, and two-layer dependency feature TLDF within the mini-program pairs, and obtain the similarity vector of the mini-program pairs;

[0101] Specifically, in step S106, the method for calculating the similarity of the layout feature LF of the mini-program pair is to traverse the layout features LF of both, regard the hash value as a string, calculate the Levenshtein ratio to compare the similarity of the two hash values. If the Levenshtein ratio is greater than a certain threshold, the number of similar hash values is incremented by one, and the number of similar hash values is divided by the maximum value of the lengths of the layout features LF of the two mini-programs. The custom function feature CFF of the mini-program is also a hash sequence, and the same similarity calculation method as the layout feature LF is used.

[0102] The method for calculating the similarity of the file dependency feature FDF of the mini-program pair is to use the Weisfeiler-Lehman algorithm to iterate the file dependency graphs of both, and calculate the similarity according to the number of similar labels in the output label set.

[0103] The method for calculating the similarity of the two-layer dependency feature TLDF of the mini-program pair is to first traverse the file dependency graphs of both. For the function call graph of the nodes, use the Weisfeiler-Lehman algorithm to calculate the similarity. If it is greater than a certain threshold, the node is put into the anchor point set, and the number of nodes in the anchor point set is divided by the maximum value of the number of nodes in the file dependency graphs of both.

[0104] Step S105: Use the similarity vectors of the mini-program pairs with pre-labeled tags of cloned and non-cloned, and use machine learning methods to construct a classifier.

[0105] Specifically, the similarity vector of the mini-program pair in step S104 is the similarity of the layout feature LF, the custom function feature CFF, the file dependency feature FDF, and the two-layer dependency feature TLDF of the two mini-programs. If two mini-programs are similar in function and layout after manual review and there are official forums or news reports of plagiarism between them, it is considered that there is a cloning relationship, and they are labeled as cloned mini-program pairs. Otherwise, they are non-cloned mini-program pairs. The input of the classifier is the similarity vector of the mini-program pair, and the output is 0 and 1. 0 represents that the mini-program pair is a non-cloned mini-program pair, and 1 represents that the mini-program pair is a cloned mini-program pair. Machine learning methods can use random forest, SVM, etc.

[0106] Step S106: According to the classifier obtained in step S105, input the similarity vector of the mini-program pair composed of the mini-program S to be detected in step S106 and each mini-program in the cloned mini-program family cluster, and perform binary classification on the mini-program pair, dividing it into cloned mini-program pairs and non-cloned mini-program pairs, so as to find the mini-program cloned with the mini-program S to be detected.< / navigator>

Claims

1. A small program cloning detection method based on complex network analysis, characterized in that, It includes the following steps: Step S101: Preprocess the to-be-detected applet S, including decompiling, deobfuscating, and extracting the main package; Step S102: According to the source code of the preprocessed to-be-detected applet S obtained in Step S101, extract statistical features SF, layout features LF, custom function features CFF, file dependency features FDF, and two-layer dependency features TLDF by analyzing the file type and the source code abstract syntax tree; Step S103: According to the statistical features SF of the to-be-detected applet S obtained in Step S102, calculate its distance from the center of the clone applet family cluster, and divide the to-be-detected applet S into the nearest clone applet family cluster; Step S104: According to the clone applet family cluster obtained in Step S103, form applet pairs by pairing the to-be-detected applet S with each applet in the clone applet family cluster. The similarity vectors of the applet pairs are the similarities of the layout features LF, custom function features CFF, file dependency features FDF, and two-layer dependency features TLDF respectively; Step S105: Use the similarity vectors of the applet pairs with pre-labeled tags of clone and non-clone, and use machine learning methods to construct a classifier; Step S106: According to the classifier obtained in Step S105, input the similarity vectors of the applet pairs formed by pairing the to-be-detected applet S with each applet in the clone applet family cluster in Step S104, and perform binary classification on the applet pairs, dividing them into clone applet pairs and non-clone applet pairs, so as to find the applets that are cloned with the to-be-detected applet S.

2. The method according to claim 1, characterized in that, The specific content of Step S101 is as follows: Step S201: Decompile the packaged file of the to-be-detected applet S according to the applet packaging rules to obtain its source code; Step S202: According to the source code of the to-be-detected applet S obtained in Step S201, extract the obfuscation pattern and match it with the existing obfuscation methods. If it can match the existing obfuscation methods, use the corresponding deobfuscation method to deobfuscate the source code; Step S203: According to the deobfuscated source code of the to-be-detected applet S obtained in Step S202, collect and organize the commonly used third-party libraries T of the applet. For the file f in the source code S of the to-be-detected applet, use the file name and file size attributes to match with the known third-party library files in T. If the file f in the S source code meets the match, filter out the third-party library files for f.

3. The method according to claim 1, wherein The statistical feature SF in Step S102 is a 241-dimensional vector. The first 3 dimensions are the number of pages of the applet, the average number of static files, and the number of developer-defined functions respectively. The 4th dimension to the 137th dimension are the call times of sourceAPI, and the 138th dimension to the 241st dimension are the call times of sinkAPI.

4. The method according to claim 1 or 3, characterized in that, The specific method for extracting the statistical feature SF in Step S102 is as follows: Step S301: According to the source code of the preprocessed to-be-detected applet S obtained in Step S101, count the number of pages of the applet from the routing components, routing functions, pages field and tabbar field of app.json; Step S302: Based on the preprocessed source code of the to-be-detected applet S obtained in Step S101, calculate the average number of static resource files in the static resource file directory by comparing the file types. The static resource file directory refers to a directory where all files are static resource files, and a static resource file refers to a file that is not a json, js, class html, or css file; Step S303: Based on the preprocessed source code of the to-be-detected applet S obtained in Step S101, count the number of functions defined by the developer himself. Functions defined by the developer himself refer to functions that are not system-defined or third-party library-defined; Step S304: Based on the preprocessed source code of the to-be-detected applet S obtained in Step S101, count the number of sourceAPI calls and sinkAPI calls. sourceAPI is an API that takes sensitive data as a return value, and sinkAPI is an API that takes sensitive data as a parameter; Step S305: Combine the data obtained in Step S301, Step S302, Step S303, and Step S304 to form a feature vector as the statistical feature SF of the to-be-detected applet S.

5. The method according to claim 1, wherein In the step S102, the layout feature LF is a hash sequence, expressed as LF(S) = <fh1, …, fh N >, where fh i represents the i-th hash value; The specific process of extracting the layout feature LF in Step S102 is as follows: Step S401: Based on the preprocessed source code of the to-be-detected applet S obtained in Step S101, parse the class html file used for page layout display, and extract the component sequence; Step S402: Shard the component sequence obtained in Step S401 using weak hashing; Step S403: Based on the sharded component sequence obtained in Step S402, calculate the hash value of the shard using strong hashing, and use the new hash sequence as its layout feature LF.

6. The method according to claim 1, characterized in that, The custom function feature CFF in the step S102 is a hash sequence, expressed as CFF(S) = <fh1, …, fh N >>, fh i represents the i-th hash value; The specific process of extracting the custom function feature CFF in Step S102 is as follows: Step S501: Based on the preprocessed source code of the to-be-detected applet S obtained in Step S101, parse the js file used for executing the function logic into an AST abstract syntax tree; Step S502: Based on the AST abstract syntax tree of the js file obtained in Step S501, extract the type sequence of the developer-defined functions. The type sequence refers to the sequence obtained by pre-order traversing the AST abstract syntax tree and sequentially taking out the type attribute values of the nodes; Step S503: Shard the component sequence obtained in Step S502 using weak hashing; Step S504: Based on the sharded component sequence obtained in Step S503, calculate the hash value of the shard using strong hashing, and use the new hash sequence as its custom function feature CFF.

7. The method according to claim 1, wherein In the step S102, the file dependency feature FDF is a file dependency graph, denoted as FDF(S) = (FileV(S), FileE(S)), where FileV(S) is the set of file dependency graph nodes of the mini-program S to be detected, and the node attribute is the file name. FileE(S) is the set of file dependency graph edges of the mini-program S to be detected, and FileE(S) = {<v i , v k >}, where v i , v k ∈ FileV(S); The specific process of extracting the file dependency feature FDF in Step S102 is as follows: Step S601: Take all js files in the source code file directory of the to-be-detected applet S as the node set of the file dependency graph; Step S602: Traverse the js files, parse them into AST abstract syntax trees, and extract the import dependency statements from the AST abstract syntax trees. If the imported file is in the node set of the file dependency graph, add this pair of files to the edge set of the file dependency graph; In step S102, the two-layer dependent feature TFDF is a two-layer graph structure. The upper layer graph is a file dependency graph, and the lower layer graph is a function call graph, denoted as FDF(S) = (FileV′(S), FileE′(S)). FileV′(S) is the set of nodes of the file dependency graph of the small program S to be detected, and the node attribute is the function call graph of this file. FileE′(S) is the set of edges of the file dependency graph of the small program S to be detected, and FileE′(S) = {<v i , v k >}, where v i , v k ∈ FileV′(S); The specific process of extracting the two-layer dependency feature TFDF in Step S102 is as follows: Step S701: Take all js files in the source code file directory of the mini program S to be detected as the set of nodes of the file dependency graph; Step S702: Traverse the js files, parse them into AST (Abstract Syntax Trees), extract the function call relationships within the files from the AST to construct a function call graph, and take the function call graph as an attribute of the js file; Step S703: According to the AST obtained in Step S702, extract the import dependency statements. If the imported file is in the set of nodes of the file dependency graph, add this pair of files to the set of edges of the file dependency graph.

8. The method according to claim 1, characterized in that In the above Step S103, first solve the center of the clone mini program family cluster, that is, calculate the average value of the statistical features SF of the mini programs within the clone mini program family cluster, calculate the distance between the statistical features SF of the mini program S to be detected and the centers of each clone mini program family cluster, and divide it into the nearest clone mini program family cluster.

9. The method according to claim 1, characterized in that In the above Step S104, the calculation method for the similarity of the layout feature LF of the mini program pair is to traverse the layout features LF of both. The hash values are regarded as strings, and the Levenshtein ratio is calculated to compare the similarity of the two hash values. If the Levenshtein ratio is greater than a certain threshold, the number of similar hash values is incremented by one, and the number of similar hash values is divided by the maximum value of the lengths of the layout features LF of the two mini programs; the custom function feature CFF of the mini program is also a hash sequence, and the same similarity calculation method as the layout feature LF is used; The calculation method for the similarity of the file dependency feature FDF of the mini program pair is to use the Weisfeiler-Lehman algorithm to iterate the file dependency graphs of both, and calculate the similarity according to the number of similar labels in the output label set; The calculation method for the similarity of the two-layer dependency feature TLDF of the mini program pair is to first traverse the file dependency graphs of both. For the function call graph of the nodes, use the Weisfeiler-Lehman algorithm to calculate the similarity. If it is greater than a certain threshold, put the node into the anchor point set, and then divide the number of nodes in the anchor point set by the maximum value of the number of nodes in the file dependency graphs of both.

10. The method according to claim 1, characterized in that In the above Step S105, the similarity vector of the mini program pair is the similarity of the layout feature LF, the similarity of the custom function feature CFF, the similarity of the file dependency feature FDF, and the similarity of the two-layer dependency feature TLDF of the two mini programs. If the two mini programs are similar in function and layout after manual review and there are official forums or news reports of plagiarism between them, it is considered that there is a cloning relationship and it is marked as a cloned mini program pair. Otherwise, it is a non-cloned mini program pair. The input of the classifier is the similarity vector of the mini program pair, and the output is 0 and 1. 0 represents that the mini program pair is a non-cloned mini program pair, and 1 represents that the mini program pair is a cloned mini program pair.

Citation Information

Patent Citations

  • Clone code detection method and device

    CN110990273A

  • Test program plagiarism detection method based on support vector machine

    CN111459788A