A method and system for identifying counterfeit applications

By comparing directory and content layer files, including source code analysis, the method improves the accuracy of identifying imposter applications on mobile devices, addressing the limitations of text-based comparisons and enhancing security.

CN114444077BActive Publication Date: 2025-07-15QIAN PANGU (SHANGHAI) INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111563436.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-20
Publication Date
2025-07-15
Estimated Expiration
2041-12-20

AI Technical Summary

Technical Problem

When identifying counterfeit applications, the prior art mainly uses text information to compare, resulting in low recognition accuracy.

Method used

Extract the result files of the target application and reference application, including directory layer files and content layer files, perform similarity analysis, combine the secure hash algorithm and the variant prefix tree to calculate the similarity, and determine whether it is a counterfeit application.

Benefits of technology

It improves the accuracy of identification of counterfeit applications, can effectively identify embedded malicious code snippets and files, and improves the accuracy of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114444077B_ABST
    Figure CN114444077B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for identifying counterfeit applications. The method includes: extracting the result files of a target application and a reference application, where the result files include directory layer files and / or content layer files; comparing the result files of the target application and the reference application to obtain the similarity between the same type of result files; and determining that the similarity is higher than a set threshold, then determining that the target application is a counterfeit application. The method and system for identifying counterfeit applications provided by the present invention determine whether a target application is a counterfeit application by comparing the directory layer files and / or content layer files of the target application and the reference application respectively, which can effectively identify the differences between the original application and the counterfeit application and improve the accuracy of identifying counterfeit applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of security technology, and in particular to a method and system for identifying counterfeit applications. Background Art

[0002] With the rapid development of mobile Internet, smart phones have become an indispensable electronic product in people's lives. The number of applications installed on mobile phones has also increased explosively. Different application functions provide mobile phone users with a good user experience.

[0003] However, with the emergence of a large number of mobile applications, the security of each application varies greatly, and there are many malicious counterfeit applications. These counterfeit applications obtain mobile phone permissions and user information by embedding malicious codes, and even have serious problems such as malicious deductions, causing huge losses to the user group. When identifying and analyzing counterfeit applications and original applications, the existing technology usually only compares and identifies them through text information, resulting in low recognition accuracy. Summary of the invention

[0004] The present invention provides a counterfeit application identification method and system to solve the defect of the prior art that only text information is compared during application identification, resulting in low identification accuracy, thereby improving the accuracy of counterfeit application identification.

[0005] The present invention provides a method for identifying counterfeit applications, comprising:

[0006] Extracting a result file of a target application and a result file of a reference application, wherein the result file includes a directory layer file and / or a content layer file; comparing the result file of the target application with the result file of the reference application to obtain a similarity between result files of the same type; and determining that the similarity is higher than a set threshold, determining that the target application is a counterfeit application.

[0007] The present invention provides a method for identifying counterfeit applications, further comprising:

[0008] If it is determined that the result file of the target application and the result file of the reference application exist in the detection platform, the result file of the target application and the result file of the reference application are extracted respectively.

[0009] If it is determined that at least one of the result file of the target application and the result file of the reference application does not exist in the detection platform, the application corresponding to the result file that does not appear in the detection platform is used as the application to be processed; it is determined whether the application to be processed is packed, and the source code file of the application to be processed is obtained based on the determination result.

[0010] If it is determined that the application to be processed is not packed, decompile the application to be processed to obtain a source code file of the application to be processed.

[0011] If it is determined that the application to be processed is shelled, then the application to be processed is unshelled to obtain an unshelled application; the unshelled application is decompiled to obtain the source code file of the application to be processed.

[0012] Compare the source code file of the target application with the source code file of the reference application to obtain the similarity between the source code files.

[0013] Create a first variant prefix tree according to the preset information in the result file of the target application, and obtain all the node numbers of the first variant prefix tree; create a second variant prefix tree according to the preset information in the result file of the reference application, and obtain all the node numbers of the second variant prefix tree; add first node numbers to the first variant prefix tree and the second variant prefix tree respectively based on a first preset strategy; add a second node number to the first variant prefix tree based on a second preset strategy; add a third node number to the second variant prefix tree based on a third preset strategy; calculate a first ratio of the first variant prefix tree based on the preset information according to the first node number, the second node number, and all the node numbers of the first variant prefix tree; calculate a second ratio of the second variant prefix tree based on the preset information according to the first node number, the third node number, and all the node numbers of the second variant prefix tree; take the ratio of the first ratio to the second ratio as the similarity.

[0014] If the result file is a directory layer file, the preset information is directory structure information; if the result file is a content layer file, the preset information is configuration file information or executable code information.

[0015] The first preset strategy is to obtain the number of sub-files with equal sha1 values in the first variant prefix tree and the second variant prefix tree based on the secure hash algorithm, and add first node numbers equal to the number of sub-files to the first variant prefix tree and the second variant prefix tree respectively.

[0016] The second preset strategy is to obtain the number of remaining sub-files with different sha1 values and the same key in the first variant prefix tree and the second variant prefix tree respectively based on a preset key, and add a second node number equal to the number of remaining sub-files with different sha1 values and the same key in the second variant prefix tree to the first variant prefix tree;

[0017] The third preset strategy is to add a third node number equal to the number of remaining sub-files with different sha1 values and the same key in the first variant prefix tree to the second variant prefix tree.

[0018] The present invention also provides a counterfeit application identification system, including:

[0019] A result file extraction unit for extracting the result file of the target application and the result file of the reference application, where the result file includes a directory layer file and / or a content layer file; a similarity acquisition unit for comparing the result file of the target application with the result file of the target application to obtain the similarity between the same type of result files; and a result output unit for determining that if the similarity is higher than a set threshold, then determining that the target application is a counterfeit application.

[0020] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any of the above-mentioned counterfeit application identification methods are implemented.

[0021] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned counterfeit application identification methods are implemented.

[0022] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned counterfeit application identification methods are implemented.

[0023] A counterfeit application identification method and system provided by the present invention first extract the result files of the target application and the reference application respectively, and then perform a comparative analysis on the directory layer file information and the content layer file information of the two result files respectively, so as to determine the similarity of the result files of the target application and the reference application. Finally, the similarity is compared with a preset threshold to determine whether the target application is a counterfeit application. Since the present invention identifies the differences between the original application and the counterfeit application by performing a comparative analysis based on the directory layer file information and the content layer file information, it is beneficial to analyze the malicious code segments and files embedded in the counterfeit application, thereby improving the accuracy of counterfeit application identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 is a schematic flowchart of the counterfeit application identification method provided by the embodiment of the present invention;

[0026] Figure 2It is a schematic flowchart of a method for identifying counterfeit applications provided by another embodiment of the present invention;

[0027] Figure 3 It is a schematic structural diagram of a counterfeit application identification system provided by an embodiment of the present invention;

[0028] Figure 4 It is a schematic structural diagram of an electronic device provided by the present invention. Detailed implementation manners

[0029] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0030] The following combines Figure 1 to describe the method for identifying counterfeit applications provided by the embodiments of the present invention, including:

[0031] Step 101: Extract the result file of the target application and the result file of the reference application, where the result file includes a directory layer file and / or a content layer file.

[0032] It can be understood that, in order to perform a similarity analysis on the target application and the reference application respectively from the directory structure information and the specific content information, it is necessary to first extract the result files containing the above two types of information from the target application and the reference application, and the result file can be obtained from a relevant application detection platform; for example, after uploading the target application A and the reference application B to the detection platform respectively, the detection platform will use a parsing algorithm to parse and package the file information contained in the application A into a result file a, and at the same time, use the same parsing algorithm to parse the file information contained in the application B and package it into a result file b; the result files a and b respectively contain the directory layer file and / or the content layer file corresponding to the applications A and B.

[0033] It should be noted that the directory layer file refers to the directory structure information in the parsing information contained in the result file, and the content layer file refers to the configuration file information and the executable code information in the parsing information contained in the result file.

[0034] Step 102: Compare the result file of the target application with the result file of the target application to obtain the similarity between the same type of result files.

[0035] Specifically, the comparison between the result file of the target application and the result file of the target application refers to: (1) the comparison of the similarity of the directory structure information in the two result files; (2) the comparison of the similarity of the configuration file information in the two result files; (3) the comparison of the similarity of the executable code information in the two result files.

[0036] Step 103: If it is determined that the similarity is higher than the set threshold, then it is determined that the target application is a counterfeit application.

[0037] It can be understood that based on the similarity value obtained in Step 102, it is compared with a preset threshold to determine whether the target file is a counterfeit application.

[0038] This embodiment provides a method for identifying counterfeit applications, which specifically analyzes the differences between the target application and the reference application from three aspects: the directory structure information, the configuration file information, and the executable code information of the target application and the reference application, which is beneficial to identifying malicious code segments and files embedded in counterfeit applications.

[0039] Optionally, if it is determined that there are the result file of the target application and the result file of the reference application in the detection platform, then the result file of the target application and the result file of the reference application are respectively extracted.

[0040] It can be understood that the detection platform can parse the file information included in the uploaded target application and reference application, and respectively package the parsed information into the result files corresponding to the two applications; for example, in this embodiment, the Janus platform is used as the detection platform. After selecting the target application on the Janus platform, a detection task is started, and then another reference application is selected on the Janus platform or uploaded from the local as the application to be compared, and then the detection task is started. If the Janus platform can directly parse the target application and the reference application through the parsing algorithm, the parsed files are respectively packaged to obtain the result files of the above two applications, and the obtained result files can be directly used for similarity comparison analysis.

[0041] It should be noted that the target application can be directly selected from the Janus platform or uploaded to the Janus platform from the local, depending on the specific application type, and the present invention does not make any restrictions.

[0042] This embodiment provides a method for obtaining the result file of the target application and the result file of the reference application, which facilitates subsequent similarity comparison analysis.

[0043] Optionally, if it is determined that at least one of the result file of the target application and the result file of the reference application does not exist in the detection platform, then the application corresponding to the result file that does not appear in the detection platform is used as the application to be processed; determine whether the application to be processed is shelled, and obtain the source code file of the application to be processed based on the determination result.

[0044] It can be understood that if the file information of an application is shelled by a manufacturer, resulting in the detection platform being unable to directly perform file parsing on the application, the required application result file cannot be obtained; at this time, a de-shelling algorithm is used to de-shell the file information of the application, and the corresponding source code file is obtained from the de-shelled application.

[0045] It should be noted that the shelling operation of an application is a process of encrypting the application, while de-shelling is the reverse process of shelling, and the de-shelling operation is a process of decrypting the shelled application; since the shelled file cannot directly obtain the source code, in order to implement the parsing of the shelled file, it is necessary to first identify the shelled file and perform de-shelling on it. Therefore, in this embodiment, the application that cannot be directly parsed through the Janus platform is downloaded into the detection sandbox to determine whether the application needs to be de-shelled.

[0046] This embodiment can de-shell the shelled applications in the comparison applications and obtain the source code files from the de-shelled applications, which provides convenience for subsequent similarity comparison and analysis.

[0047] Optionally, if it is determined that the application to be processed is not shelled, then the application to be processed is decompiled to obtain the source code file of the application to be processed.

[0048] It can be understood that if this embodiment uses the detection sandbox to determine that the application to be processed is not shelled, the source code file is obtained by decompiling the application, and the encrypted information is no longer included in the obtained source code file, which improves the purity of the information required in the source code file.

[0049] This embodiment provides a method for obtaining the source code file of an unshelled application when it is determined that the application to be processed is not shelled, which provides convenience for subsequent similarity comparison and analysis.

[0050] Optionally, if it is determined that the application to be processed is shelled, then the application to be processed is de-shelled to obtain the de-shelled application; the de-shelled application is decompiled to obtain the source code file of the application to be processed.

[0051] It can be understood that if this embodiment uses the detection sandbox to determine that the application to be processed is shelled, it is necessary to use the de-shelling program supported by the detection sandbox to de-shell the shelled application, and then decompile the de-shelled application to obtain the source code file.

[0052] This embodiment provides a method for obtaining the source code file of an application to be processed when it is detected that the application is shelled, which facilitates subsequent similarity comparison and analysis.

[0053] Optionally, compare the source code file of the target application with the source code file of the reference application to obtain the similarity between the source code files.

[0054] It can be understood that after the target application or the reference application in this embodiment passes the shell detection, the corresponding source code file is obtained through decompilation and encoding, and then similarity analysis is performed based on the file information included in the two source code files. Specifically, it includes: (1) comparing the similarity of the directory structure information in the two source code files; (2) comparing the similarity of the configuration file information in the two source code files; (3) comparing the similarity of the executable code information in the two source code files.

[0055] It should be noted that in order to improve the efficiency of similarity analysis, this embodiment can package the two obtained source code files into the result file of the target application and the result file of the reference application provided in the above embodiment respectively. The result file also includes the directory structure information, configuration file information and executable code information of the corresponding application.

[0056] This embodiment can obtain the source code file from the shell-detected application and perform similarity comparison, improving the purity of the file types included in the source code file.

[0057] Optionally, create a first variant prefix tree according to the preset information in the result file of the target application, and obtain all the node numbers of the first variant prefix tree; create a second variant prefix tree according to the preset information in the result file of the reference application, and obtain all the node numbers of the second variant prefix tree; add the first node numbers to the first variant prefix tree and the second variant prefix tree respectively based on the first preset strategy; add the second node number to the first variant prefix tree based on the second preset strategy; add the third node number to the second variant prefix tree based on the third preset strategy; calculate the first ratio of the first variant prefix tree based on the preset information according to the first node number, the second node number and all the node numbers of the first variant prefix tree; calculate the second ratio of the second variant prefix tree based on the preset information according to the first node number, the third node number and all the node numbers of the second variant prefix tree; use the ratio of the first ratio to the second ratio as the similarity.

[0058] Specifically, to facilitate the similarity comparison based on the preset information included in the result file of the target application and the result file of the reference application, using the preset information as the partitioning attribute, the two result files can be respectively created as the first variant prefix tree and the second variant prefix tree. Among them, the number of nodes in the first variant prefix tree is the number of files containing the preset information in the result file of the target application, and the number of nodes in the second variant prefix tree is the number of files containing the preset information in the result file of the control application. Then, according to the first preset strategy, the number of nodes corresponding to the file information that exists in the second variant prefix tree but does not exist in the first variant prefix tree is added to the nodes in the first variant prefix tree. According to the second preset strategy, the number of nodes corresponding to the file information that exists in the first variant prefix tree but does not exist in the second variant prefix tree is added to the nodes in the second variant prefix tree. Next, calculate the first ratio of the number of nodes corresponding to the same information in the first variant prefix tree and the second variant prefix tree to the total number of nodes in the first variant prefix tree after adding nodes. At the same time, calculate the second ratio of the number of nodes corresponding to the same information in the second variant prefix tree and the first variant prefix tree to the total number of nodes in the second variant prefix tree after adding nodes. Finally, take the ratio of the first ratio to the second ratio as the similarity value of the above two result files. For example, in an embodiment, the number of nodes containing the preset information in the result file of the target application and the result file of the reference application are M and N respectively. Based on the first preset strategy, the first number of nodes x is added to the first variant prefix tree and the second variant prefix tree respectively. Based on the second preset strategy, the second number of nodes y is added to the first variant prefix tree. Based on the third preset strategy, the third number of nodes z is added to the second variant prefix tree, where M and N are natural numbers greater than 2, and x, y, and z are natural numbers less than M and N. Then the way to obtain the first ratio is P(M) = ((x + z)); / The way to obtain the second ratio P(N) = (x + z) / (x + z + N), and the calculation method of the similarity value of the two result files is P(M) / P(N).

[0059] This embodiment provides a similarity calculation method based on result files, which is used to quantitatively describe the similarity degree between the target application and the reference application.

[0060] Optionally, if the result file is a directory layer file, the preset information is directory structure information; if the result file is a content layer file, the preset information is configuration file information or executable code information.

[0061] It can be understood that the preset information can be the directory structure information in the result file, or the configuration file information in the result file, or the executable code information in the result file. In this embodiment, the total number of nodes for creating the variant prefix tree can be respectively based on the number of file information of the three types. Then, based on the similarity calculation method of the previous embodiment, the application similarity corresponding to different preset information is calculated.

[0062] In this embodiment, the preset information may be the directory structure information, configuration file information, and executable code information in the result file. The similarity values of two comparison applications corresponding to the directory layer file and the content layer file of the result file can be calculated according to different preset information.

[0063] Optionally, the first preset strategy is to obtain the number of sub-files with equal sha1 values in the first variant prefix tree and the second variant prefix tree based on the secure hash algorithm, and add the same number of first nodes as the number of sub-files in the first variant prefix tree and the second variant prefix tree respectively.

[0064] Specifically, this embodiment provides a specific implementation manner for determining the first number of nodes by the first preset strategy. Among them, the secure hash algorithm is a secure algorithm, which mainly determines whether two comparison files are the same by calculating whether the sha1 values of the two comparison files are equal.

[0065] This embodiment can determine the number of nodes with the same sha1 value in the first variant prefix tree and the second variant prefix tree based on the preset information, which is convenient for calculating the similarity values of the two applications subsequently.

[0066] Optionally, the second preset strategy is to obtain the number of remaining sub-files with different sha1 values and the same key in the first variant prefix tree and the second variant prefix tree based on the preset key, and add the same number of second nodes as the number of remaining sub-files with different sha1 values and the same key in the second variant prefix tree to the first variant prefix tree;

[0067] Specifically, this embodiment provides a specific implementation manner for determining the second number of nodes by the second preset strategy, and adds the determined second number of nodes to the first variant prefix tree; among them, the response parameters with the same file name and the same number of levels are used as the key, and the key is often used for the verification code of the software.

[0068] This embodiment can determine the second number of nodes added to the first variant prefix tree, which is convenient for calculating the first ratio subsequently.

[0069] Optionally, the third preset strategy is to add the same number of third nodes as the number of remaining sub-files with different sha1 values and the same key in the first variant prefix tree to the second variant prefix tree.

[0070] Specifically, this embodiment provides a specific implementation manner for determining the third number of nodes by the third preset strategy, and adds the determined third number of nodes to the second variant prefix tree.

[0071] This embodiment can determine the number of third nodes added to the second variant prefix tree, which facilitates the subsequent calculation of the second ratio.

[0072] Based on the above embodiments, combined with Figure 2 A specific description of the implementation process of the method provided by the present invention is given.

[0073] In this embodiment, first directly select the application to be detected 1 from the Janus platform, and upload the application to be detected 2 from the local to the Janus platform. Since there are no result files for the application to be detected 1 and the application to be detected 2 on the Janus platform, the two cannot be directly compared for similarity; at this time, download the application to be detected 1 and the application to be detected 2 to the detection sandbox for shell detection respectively. It is detected that the application to be detected 1 has been shelled, while the application to be detected 2 has not been shelled. Then, use the built-in de-shelling algorithm in the detection sandbox to de-shell the application to be detected 1, and then perform decompilation on the de-shelled application to be detected 1 to obtain the source code file 1. Then, perform parsing operations on the source code file 1 and store the parsed results as the result file 1. At the same time, perform decompilation on the application to be detected 2 to obtain the source code file 2, then perform parsing operations on the source code file 2 and store the parsed results as the result file 2. Finally, use the same similarity calculation method as in the above embodiments to perform similarity comparison and analysis on the result file 1 and the result file 2, so as to determine the similarity comparison of the directory structure information, the similarity comparison of the configuration file information, and the similarity of the executable code information in the two result files. After comparing with the preset threshold, finally determine the counterfeit application.

[0074] A method for identifying counterfeit applications provided by this embodiment first extracts the result files of the target application and the reference application respectively, and then performs comparative analysis on the directory-level file information and content-level file information of the two result files respectively, so as to determine the similarity of the result files of the target application and the reference application. Finally, compare the similarity with a preset threshold to determine whether the target application is a counterfeit application. Since the present invention identifies the differences between the original application and the counterfeit application by performing comparative analysis on the directory-level file information and content-level file information, it is beneficial to analyze the malicious code segments and files embedded in the counterfeit application, thereby improving the accuracy of counterfeit application identification.

[0075] Combined with Figure 3 A counterfeit application identification system provided by an embodiment of the present invention is described. A data set partitioning device described below can be correspondingly referred to the counterfeit application identification system described above.

[0076] A counterfeit application identification system provided by the present invention includes:

[0077] The result file extraction unit 301 is configured to extract the result file of the target application and the result file of the reference application, where the result file includes a directory layer file and / or a content layer file.

[0078] The similarity acquisition unit 302 is configured to compare the result file of the target application with the result file of the target application to obtain the similarity between the same type of result files.

[0079] The result output unit 303 is configured to determine that if the similarity is higher than a set threshold, then determine that the target application is a counterfeit application.

[0080] The system in this embodiment extracts the result file of the target application and the result file of the reference application through the result file extraction unit 301, then obtains the similarity between the same type of result files by comparing the result file of the target application with the result file of the target application through the similarity acquisition unit 302, and finally compares the similarity value with the set threshold through the result output unit 303, so as to determine whether the target application is a counterfeit application. The system in this embodiment can analyze the differences between the target application and the reference application from three aspects: the directory structure information, the configuration file information, and the executable code information of the target application and the reference application, which is beneficial to identifying malicious code segments and files embedded in the counterfeit application.

[0081] Figure 4 An entity structure diagram of an electronic device is exemplified, as Figure 4 shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 complete mutual communication through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute a method for identifying a counterfeit application, and the method includes: extracting the result file of the target application and the result file of the reference application, where the result file includes a directory layer file and / or a content layer file; comparing the result file of the target application with the result file of the reference application to obtain the similarity between the same type of result files; determining that if the similarity is higher than a set threshold, then determine that the target application is a counterfeit application.

[0082] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0083] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to implement a method for identifying counterfeit applications provided by the above-mentioned various methods. The method includes: extracting the result file of the target application and the result file of the reference application, where the result file includes a directory layer file and / or a content layer file; comparing the result file of the target application and the result file of the reference application to obtain the similarity between the same type of result files; and determining that the similarity is higher than a set threshold, then determining that the target application is a counterfeit application.

[0084] The present invention also provides a computer program product. The computer program product includes a computer program, and the computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a method for identifying counterfeit applications provided by the above-mentioned various methods. The method includes: extracting the result file of the target application and the result file of the reference application, where the result file includes a directory layer file and / or a content layer file; comparing the result file of the target application and the result file of the reference application to obtain the similarity between the same type of result files; and determining that the similarity is higher than a set threshold, then determining that the target application is a counterfeit application.

[0085] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0086] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A method for identifying counterfeit applications, characterized in that, Including: Extracting the result file of the target application and the result file of the reference application, where the result file includes a directory layer file and / or a content layer file; Comparing the result file of the target application and the result file of the reference application to obtain the similarity between the same type of result files; Determining that the similarity is higher than a set threshold, then determining that the target application is a counterfeit application; The comparing the result file of the target application and the result file of the reference application to obtain the similarity between the same type of result files specifically includes: Creating a first variant prefix tree according to the preset information in the result file of the target application, and obtaining all the node numbers of the first variant prefix tree; Creating a second variant prefix tree according to the preset information in the result file of the reference application, and obtaining all the node numbers of the second variant prefix tree; Adding a first number of nodes to the first variant prefix tree and the second variant prefix tree respectively based on a first preset strategy; Adding a second number of nodes to the first variant prefix tree based on a second preset strategy; Adding a third number of nodes to the second variant prefix tree based on a third preset strategy; Calculating a first ratio of the first variant prefix tree based on the preset information according to the first number of nodes, the second number of nodes, and all the node numbers of the first variant prefix tree; Calculating a second ratio of the second variant prefix tree based on the preset information according to the first number of nodes, the third number of nodes, and all the node numbers of the second variant prefix tree; Taking the ratio of the first ratio to the second ratio as the similarity.

2. The method for identifying a counterfeit application according to claim 1, wherein The extracting the result file of the target application and the result file of the reference application specifically includes: Determining that there are the result file of the target application and the result file of the reference application in the detection platform, then respectively extracting the result file of the target application and the result file of the reference application.

3. The counterfeit application identification method according to claim 1, wherein The extracting the result file of the target application and the result file of the reference application specifically includes: Determining that at least one of the result file of the target application and the result file of the reference application does not exist in the detection platform, then taking the application corresponding to the result file that does not appear in the detection platform as the application to be processed; Judging whether the application to be processed is shelled, and obtaining the source code file of the application to be processed based on the judgment result.

4. The counterfeit application identification method according to claim 3, characterized in that The obtaining the source code file of the application to be processed based on the judgment result specifically includes: Determining that the application to be processed is not shelled, then decompiling the application to be processed to obtain the source code file of the application to be processed.

5. The method for identifying a counterfeit application according to claim 3, wherein The obtaining the source code file of the application to be processed based on the judgment result specifically includes: Determining that the application to be processed is shelled, then unshelling the application to be processed to obtain an unshelled application; Decompiling the unshelled application to obtain the source code file of the application to be processed.

6. The anti-counterfeiting application identification method according to any one of claims 3-5, characterized in that, The comparing the result file of the target application and the result file of the reference application to obtain the similarity between the same type of result files specifically includes: Comparing the source code file of the target application and the source code file of the reference application to obtain the similarity between the source code files.

7. The method for identifying counterfeit applications according to claim 1, wherein if the result file is a directory-level file, the preset information is directory structure information; if the result file is a content-level file, the preset information is configuration file information or executable code information.

8. The method for identifying counterfeit applications according to claim 1, wherein the first preset strategy is to obtain the number of sub-files with equal sha1 values in the first variant prefix tree and the second variant prefix tree based on the secure hash algorithm, and add the same number of first nodes as the number of sub-files in the first variant prefix tree and the second variant prefix tree respectively.

9. The method for identifying counterfeit applications according to claim 1, wherein the second preset strategy is to obtain the number of remaining sub-files with different sha1 values and the same key in the first variant prefix tree and the second variant prefix tree based on a preset key, and add the same number of second nodes as the number of remaining sub-files with different sha1 values and the same key in the second variant prefix tree in the first variant prefix tree.

10. The method for identifying counterfeit applications according to claim 1, wherein the third preset strategy is to add the same number of third nodes as the number of remaining sub-files with different sha1 values and the same key in the first variant prefix tree in the second variant prefix tree.

11. An imitation application recognition system, characterized in that, It includes: a result file extraction unit for extracting the result file of the target application and the result file of the reference application, where the result file includes a directory-level file and / or a content-level file; a similarity acquisition unit for comparing the result file of the target application and the result file of the reference application to obtain the similarity between the result files of the same type; a result output unit for determining that if the similarity is higher than a set threshold, it is determined that the target application is a counterfeit application; The comparison of the result file of the target application and the result file of the reference application to obtain the similarity between the result files of the same type specifically includes: creating a first variant prefix tree according to the preset information in the result file of the target application, and obtaining all the node numbers of the first variant prefix tree; creating a second variant prefix tree according to the preset information in the result file of the reference application, and obtaining all the node numbers of the second variant prefix tree; adding the first node numbers to the first variant prefix tree and the second variant prefix tree respectively based on the first preset strategy; adding the second node numbers to the first variant prefix tree based on the second preset strategy; adding the third node numbers to the second variant prefix tree based on the third preset strategy; calculating a first ratio of the first variant prefix tree based on the preset information according to the first node numbers, the second node numbers and all the node numbers of the first variant prefix tree; calculating a second ratio of the second variant prefix tree based on the preset information according to the first node numbers, the third node numbers and all the node numbers of the second variant prefix tree; taking the ratio of the first ratio to the second ratio as the similarity.

12. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of the counterfeit application recognition method according to any one of claims 1 to 10 are implemented.

13. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the counterfeit application recognition method according to any one of claims 1 to 10 are implemented.

14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the counterfeit application recognition method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Method and device for detecting safety performance of application program

    CN104123493A