Application Reinforcement Detection Method, Device, Readable Medium and Electronic Device
By parsing and obfuscating the application's source code file to determine whether it is a random string, the problem of low stability of traditional machine learning detection methods is solved, and more accurate and stable reinforcement detection is achieved.
Patent Information
- Application Number
- CN202111258495.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-10-27
AI Technical Summary
When traditional machine learning detection methods are used in the prior art, the stability of application reinforcement detection is not high, resulting in inaccurate detection results.
By obtaining multiple source code files of the application, each source code file is parsed and obfuscated, and it is determined whether it is a random string, and then whether the application has undergone code hardening.
Improve the accuracy and stability of application reinforcement detection, ensuring that the judgment of code reinforcement is more comprehensive and accurate.
Smart Images

Figure CN114281669B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and particularly relates to an application program hardening detection method, device, readable medium, and electronic device. Background Technique
[0002] Having good security is one of the important conditions for an application program to run stably. Code hardening is a means to improve application security. Code obfuscation is a commonly used code hardening method. Code obfuscation is the act of converting the code of a computer program into a form that is functionally equivalent but difficult to read and understand. Therefore, to detect whether an application program is secure, it is usually determined by detecting whether the application program has undergone code obfuscation. Existing code obfuscation detection methods usually use traditional machine learning detection methods: first, extract the text features of the code files of the application program, and then use traditional binary classification models such as support vector machines to classify the text features, and finally determine whether the application program is obfuscated and hardened. This detection method depends on the feature extraction process. When the extracted features are inappropriate, the detection results are inaccurate and the stability needs to be improved.
[0003] It should be noted that the information disclosed in the above background section is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The purpose of this application is to provide an application program hardening detection method, device, readable medium, and electronic device to solve the problem of low stability of using traditional machine learning detection methods in related technologies.
[0005] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.
[0006] According to one aspect of the embodiments of this application, an application program hardening detection method is provided, including:
[0007] Obtain multiple source code files of the application program;
[0008] Parse each of the source code files to obtain multiple strings corresponding to the source code files;
[0009] Perform obfuscation detection on the source code files by determining whether each string corresponding to the source code file is a random string, and obtain the obfuscation detection result of the source code file. The obfuscation detection result is an obfuscated code file or an unobfuscated code file;
[0010] Determine whether the application is fortified according to the obfuscation detection results of all source code files.
[0011] According to one aspect of the embodiments of the present application, there is provided an application fortification detection device, including:
[0012] A source code acquisition module, configured to acquire multiple source code files of an application;
[0013] A code parsing module, configured to parse each of the source code files respectively to obtain multiple strings corresponding to the source code files;
[0014] An obfuscation detection module, configured to perform obfuscation detection on each string corresponding to the source code file to obtain the obfuscation detection result of the source code file, where the obfuscation detection result is an obfuscated code file or an unobfuscated code file;
[0015] A code fortification determination module, configured to determine whether the application is fortified according to the obfuscation detection results of all source code files.
[0016] In an embodiment of the present application, the string includes a variable string and a common string; the obfuscation detection module includes:
[0017] A variable string detection unit, configured to perform obfuscation detection on the variable string through a simple character recognition algorithm and a random string recognition model, and obtain a variable obfuscation ratio according to the obfuscation detection results of all variable strings, where the variable obfuscation ratio is the ratio of the random strings in the variable string;
[0018] A common string detection unit, configured to perform obfuscation detection on the common string through the random string recognition model, and obtain a common string obfuscation ratio according to the obfuscation detection results of all common strings, where the common string obfuscation ratio is the ratio of the random strings in the common string;
[0019] An obfuscation result determination unit, configured to determine the obfuscation detection result of the source code file according to the variable obfuscation ratio and the common string obfuscation ratio.
[0020] In an embodiment of the present application, the variable string detection unit is specifically configured to:
[0021] Determine whether the variable string is a simple string through the simple character recognition algorithm; if the variable string is a simple string, determine that the variable string is a random string; if the variable string is not a simple string, determine whether the variable string is a random string through the random string recognition model; determine the variable obfuscation ratio according to the ratio of the number of variable strings determined to be random strings to the total number of variable strings.
[0022] In one embodiment of the present application, the obfuscation result determination unit is specifically configured to:
[0023] When the variable obfuscation ratio is greater than a first preset character threshold and the ordinary string obfuscation ratio is greater than a second preset character threshold, determine that the obfuscation detection result of the source code file is an obfuscated code file; when the variable obfuscation ratio is less than the first preset character threshold or the ordinary string obfuscation ratio is less than the second preset character threshold, determine that the obfuscation detection result of the source code file is an unobfuscated code file.
[0024] In one embodiment of the present application, the device further includes:
[0025] A sample set acquisition module, configured to acquire a sample data set, where the sample data set includes sample strings;
[0026] A training set generation module, configured to perform random transformation and splicing processing on the sample strings in the sample data set to form a training sample set;
[0027] A model training module, configured to train a neural network model based on the training sample set to obtain the random string recognition model.
[0028] In one embodiment of the present application, the sample strings include Chinese pinyin, English words, and delimiters, and the training sample set includes positive sample data; the training set generation module includes:
[0029] A first random transformation unit, configured to perform random case transformation on the Chinese pinyin to obtain extended Chinese pinyin; perform random case transformation on the English words to obtain extended English words;
[0030] A first splicing unit, configured to splice the extended Chinese pinyin, the extended English words, and the delimiters to form the positive sample data of the training sample set.
[0031] In one embodiment of the present application, the first random transformation unit is configured to perform at least one of the following operations:
[0032] Randomly select n extended Chinese pinyin and n + 1 delimiters, and perform alternating splicing processing on the n extended Chinese pinyin and the n + 1 delimiters;
[0033] Randomly select n extended English words and n + 1 delimiters, and perform alternating splicing processing on the n extended English words and the n + 1 delimiters;
[0034] Randomly select m extended Chinese pinyins, n - m extended English words, and n + 1 delimiters, and perform an alternating splicing process on the n extended Chinese pinyins, extended English words, and n + 1 delimiters; m < n, and n and m are natural numbers greater than 0.
[0035] In an embodiment of the present application, the sample string includes uppercase English characters, lowercase English characters, and delimiters, and the training sample set includes negative sample data; the training set generation module includes:
[0036] A second random transformation unit, configured to generate n character units according to the uppercase English characters and the lowercase English characters, where one character unit is composed of a random number of English characters selected from the uppercase English characters and the lowercase English characters;
[0037] A second splicing unit, configured to perform an alternating splicing process on the n character units and randomly selected n + 1 delimiters to form the negative sample data of the training sample set.
[0038] In an embodiment of the present application, the model training module includes:
[0039] A data preprocessing unit, configured to preprocess the training sample set to form an input data set for the neural network model; the preprocessing includes: determining the maximum length of the sample data in the training sample set, where the length of the sample data refers to the number of characters included in the sample data; splicing the sample data with a preset character if the length is less than the maximum length to form input data with a length reaching the maximum length;
[0040] A model training unit, configured to train the neural network model based on the input data set to obtain the random string recognition model.
[0041] In an embodiment of the present application, after parsing the source code file, interface call data in the source code file is also obtained; the apparatus further includes:
[0042] A logic detection module, configured to perform logic detection on the source code file according to the interface call data of the source code file to obtain a logic detection result of the source code file, where the logic detection result includes a logic - hidden code file and a logic - unhidden code file;
[0043] Correspondingly, the code hardening determination module is configured to:
[0044] Determine whether the application program is hardened according to the obfuscation detection results of all source code files and the logic detection results of all source code files.
[0045] In one embodiment of the present application, the interface call data includes explicit interface call data and implicit interface call data; specifically, the logic detection module is configured to:
[0046] Determine the proportion of implicit interface calls in the source code file based on the explicit interface call data and implicit interface call data of the source code file; when the proportion of implicit interface calls is less than a preset interface threshold, determine that the logic detection result of the source code file is a logically unhidden code file; when the proportion of implicit interface calls is greater than the preset interface threshold, determine that the logic detection result of the source code file is a logically hidden code file.
[0047] In one embodiment of the present application, the code hardening determination module is specifically configured to:
[0048] Determine the proportion of obfuscated code files based on the number of source code files that are obfuscated code files and the total number of source code files; determine the proportion of logically hidden code files based on the number of source code files whose logic detection result is a logically hidden code file and the total number of source code files; when the proportion of obfuscated code files is greater than a first preset file threshold and the proportion of logically hidden code files is greater than a second preset file threshold, determine that the application has been hardened; when the proportion of obfuscated code files is less than the first preset file threshold or the proportion of logically hidden code files is less than the second preset file threshold, determine that the application has not been hardened.
[0049] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the application hardening detection method in the above technical solution.
[0050] According to one aspect of the embodiments of the present application, there is provided an electronic device, which includes: a processor; and a memory for storing executable instructions of the processor; wherein, when the processor executes the executable instructions, the electronic device executes the application hardening detection method in the above technical solution.
[0051] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the application hardening detection method in the above technical solution.
[0052] In the technical solution provided by the embodiments of the present application, by obtaining multiple source code files of an application program, then parsing, obfuscation detection, and logic detection are respectively performed on each source code file, and finally, it is determined whether the application program is fortified according to the obfuscation detection results and logic detection results of all source code files; it realizes judging whether the application program is fortified from two aspects of obfuscation detection and logic detection, the judgment of code fortification is more comprehensive, and the judgment result is more accurate. In addition, the application program is divided into multiple source code files for obfuscation detection and logic detection, that is, the detection granularity of the fortification detection is refined, further improving the accuracy of the judgment result.
[0053] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0055] Figure 1 Schematically shows an exemplary system architecture block diagram applying the technical solution of the present application.
[0056] Figure 2 Schematically shows a flowchart of an application program fortification detection method provided by an embodiment of the present application.
[0057] Figure 3 Schematically shows a flowchart of obfuscation detection for a source code file provided by an embodiment of the present application.
[0058] Figure 4 Schematically shows a flowchart of generating a random string recognition model provided by the present application.
[0059] Figure 5 Schematically shows a flowchart of generating positive sample data provided by the present application.
[0060] Figure 6 Schematically shows a flowchart of generating negative sample data provided by the present application.
[0061] Figure 7 Schematically shows the structure diagram of the CharCNN model provided by the present application.
[0062] Figure 8 Shows a flowchart of applet obfuscation detection provided by an embodiment of the present application.
[0063] Figure 9 The flowchart of the mini-program logic detection provided by an embodiment of the present application is shown.
[0064] Figure 10 The flowchart of the mini-program reinforcement detection method provided by an embodiment of the present application is shown.
[0065] Figure 11 The structural block diagram of the application program reinforcement detection device provided by the embodiment of the present application is schematically shown.
[0066] Figure 12 The computer system structural block diagram of the electronic device for implementing the embodiment of the present application is schematically shown. Detailed implementation manners
[0067] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0068] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0069] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0070] The flowcharts shown in the drawings are only exemplary illustrations and do not necessarily include all the content and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0071] Figure 1 The exemplary system architecture block diagram applying the technical solution of the present application is schematically shown.
[0072] Such as Figure 1As shown in the figure, the system architecture 100 may include a terminal device 110, a network 120, and a server 130. The terminal device 110 may include, but is not limited to, a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc. The server 130 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The network 120 may be a communication medium of various connection types capable of providing a communication link between the terminal device 110 and the server 130. For example, it may be a wired communication link or a wireless communication link.
[0073] According to the implementation requirements, the system architecture in the embodiments of the present application may have any number of terminal devices, networks, and servers. For example, the server 130 may be a server group composed of multiple server devices. In addition, the technical solutions provided in the embodiments of the present application may be applied to the terminal device 110, or may be applied to the server 130, or may be jointly implemented by the terminal device 110 and the server 130. The present application makes no special limitation thereto.
[0074] In an embodiment of the present application, the application program hardening detection method provided in the embodiments of the present application is executed by the server 130. Correspondingly, the application program hardening detection device is generally disposed in the server 130. Specifically, the server 130 obtains a plurality of source code files of the application program; then parses each source code file respectively to obtain a plurality of strings corresponding to the source code file. Then the server 130 performs obfuscation detection on the source code file by determining whether each string corresponding to the source code file is a random string, and obtains the obfuscation detection result of the source code file. The obfuscation detection result is an obfuscated code file or an unobfuscated code file. Finally, the server 130 determines whether the application program is hardened according to the obfuscation detection results of all source code files. For example, when the number of obfuscated code files reaches a preset threshold, it is determined that the application program has been hardened.
[0075] In an embodiment of the present application, after parsing each source code file, logical call data in the source code file may also be obtained. The server 130 may perform logical detection on the source code file according to the interface call data of the source code file, and obtain the logical detection result of the source code file. The logical detection result is a logically hidden code file or a logically unhidden code file. Finally, the server 130 determines whether the application program is hardened according to the obfuscation detection results of all source code files and the logical detection results of all source code files. For example, when the number of obfuscated code files and the number of logically hidden code files both reach a certain threshold, it is determined that the application program has been hardened.
[0076] In an embodiment of the present application, artificial intelligence technology can be used to perform obfuscation detection or logic detection on source code files. Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0077] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0078] In an embodiment of the present application, after the server 130 determines whether an application is fortified, it can confirm the result and feedback it to the terminal device 110 through the network 120. Then, the terminal device 110 can display the detection result of whether the application is fortified to the user.
[0079] In an embodiment of the present application, those skilled in the art can easily understand that the application fortification detection method provided by the embodiments of the present application can also be executed by the terminal device 110. Correspondingly, the application fortification detection device can also be set in the terminal device 110, and no special limitation is made in this exemplary embodiment. For example, in an exemplary embodiment, the terminal device 110 obtains multiple source code files of an application, performs parsing on the source code files, and then performs obfuscation detection and logic detection respectively. Finally, it determines whether the application is fortified according to the obfuscation detection results of all source code files and the logic detection results of all source code files.
[0080] The following will make a detailed description of the application fortification detection method provided by the present application in combination with specific embodiments.
[0081] Figure 2 Schematically shows a flowchart of an application fortification detection method provided by an embodiment of the present application. This method can be executed by a terminal device, such as Figure 1 the terminal device 110 shown; this method can also be executed by a server, such as Figure 1 the server 130 shown. As Figure 2As shown in the figure, the application program hardening detection method provided by the embodiment of the present application includes steps 210 to 240, which are specifically as follows:
[0082] Step 210, obtain multiple source code files of the application program.
[0083] Specifically, an application program refers to a computer program for completing a certain or multiple specific tasks, and this computer program is obtained by compiling source code; at the same time, the application program runs in user mode, can interact with users, and has a visible user interface. A source code file refers to a code file composed of source code. Source code is a human-readable text file written according to certain programming language specifications, that is, the original code written by developers, usually written in high-level languages such as Java language and Javascript language. Generally, an application program is composed of multiple source code files.
[0084] In an embodiment of the present application, the process of obtaining multiple source code files of the application program specifically includes: obtaining the executable code file of the application program; performing decompilation processing on the executable code file to obtain multiple source code files.
[0085] Specifically, an executable code file refers to a code file composed of computer-executable program code (that is, the computer program described above), and computer-executable program code usually refers to binary code. The process of converting source code into executable code is called compilation. Then, through decompilation processing, the executable code can be restored to source code.
[0086] Step 220, parse each source code file respectively to obtain multiple strings corresponding to the source code file.
[0087] Specifically, after obtaining multiple source code files, parsing a source code file can obtain multiple strings corresponding to the source code file. The strings of the source code file refer to the strings composed of code characters in the source code file. For example, if the source code file includes: var nameStr = "jack", then both "nameStr" and "jack" are strings, where "var" represents the variable declaration method and does not belong to the strings described in the present application.
[0088] In one embodiment of the present application, an Abstract Syntax Code (AST) parsing method is used to parse the source code file. An AST is a tree-like representation of the abstract syntax structure of the source code, and each node on the tree represents a structure in the source code. "Abstract" means that the AST does not represent every detail of the actual syntax. For example, nested parentheses are implied in the tree structure and are not presented in the form of nodes, for instance. The AST does not depend on the grammar of the source language, that is, the context-free grammar used in the syntax analysis phase. Exemplarily, the esprima tool can be used to perform AST parsing on the source code file.
[0089] Step 230: Perform obfuscation detection on the source code file by determining whether each string corresponding to the source code file is a random string, and obtain the obfuscation detection result of the source code file. The obfuscation detection result is an obfuscated code file or an unobfuscated code file.
[0090] Specifically, performing obfuscation detection on the source code file is to detect whether the code in the source code file has been obfuscated. The obfuscation of the code is the act of converting the normally readable code into a functionally equivalent but difficult-to-read and understand form. If the code in the source code file has not been obfuscated, it means that the source code file is an unobfuscated code file; if the code in the source code file has been obfuscated, it means that the source code file is an obfuscated code file.
[0091] In an embodiment of the present application, obfuscation detection of the source code file is achieved by determining whether each string corresponding to the source code file is a random string. A random string is a randomly generated string that usually has no specific meaning and is difficult to read and understand. When performing code obfuscation operations, normal code strings are usually changed to random strings. When the code is leaked, others will see random strings instead of normal code strings, so that the correct code content cannot be obtained. Therefore, the operation of changing normal code strings to random strings can prevent the leakage of correct code.
[0092] In one embodiment of the present application, it is possible to determine whether the source code file is an obfuscated code file based on the number of random strings in the source code file. For example, when the number of random strings in the source code file reaches a certain value, it is considered that the source code file is an obfuscated code file; or, when the number of random strings in the source code file is greater than the number of non-random strings, it is considered that the source code file is an obfuscated code file.
[0093] In one embodiment of the present application, the strings in the source code file include two types: variable strings and ordinary strings. A variable string refers to a string formed by variable codes in the source code, and an ordinary string is a string other than the variable string. For example, if the source code file includes: var nameStr = "jack", then "nameStr" is a variable string and "jack" is an ordinary string.
[0094] In one embodiment of the present application, as Figure 3 shown, the steps of performing obfuscation detection on the source code file by determining whether each string corresponding to the source code file is a random string include steps 310 to 330, specifically as follows:
[0095] Step 310: Perform obfuscation detection on the variable strings through a simple character recognition algorithm and a random string recognition model, and obtain a variable obfuscation ratio based on the obfuscation detection results of all variable strings. The variable obfuscation ratio is the ratio of the random strings in the variable strings.
[0096] Specifically, performing obfuscation detection on a string is to detect whether the string is a random string. In the embodiment of the present application, there are two methods to determine whether a variable string is a random string. The simple character recognition algorithm is used to identify whether a variable string is a simple string. A simple string refers to a string with a simple structure and a short length. The length of a string refers to the number of characters contained in the string. For example, a string with a length less than or equal to 2 can be regarded as a simple string. For example, a single character a or a double character ab, etc. can be determined as a simple string. The random string recognition model is used to identify whether a variable string is a random string. The random string recognition model is a neural network model pre-trained according to training samples. Finally, the ratio of the number of variable strings determined to be random strings to the total number of variable strings is determined to obtain the variable obfuscation ratio.
[0097] In one embodiment of the present application, step 310 is specifically: determining whether a variable string is a simple string through a simple character recognition algorithm; if the variable string is a simple string, determining that the variable string is a random string; if the variable string is not a simple string, determining whether the variable string is a random string through a random string recognition model; and determining the variable obfuscation ratio according to the ratio of the number of variable strings determined to be random strings to the total number of variable strings.
[0098] Specifically, first, simple strings in the variable string are screened out through a simple character recognition algorithm. Since simple strings usually have no actual meaning, the simple strings are directly determined as random strings. When it is determined by the simple character recognition algorithm that the variable string is not a simple string, at this time, it is further determined whether the variable string is a random string through a random string recognition model. The random string recognition model can recognize whether relatively complex strings are random strings. Through the combination of the simple character recognition algorithm and the random string recognition model, the confusion detection of the variable string can be completed.
[0099] After performing confusion detection on all variable strings of a source code file, count the number of variable strings determined as random strings, and then use the ratio of this number to the total number of variable strings as the variable confusion ratio α corresponding to the source code file.
[0100] Step 320, perform confusion detection on ordinary strings through a random string recognition model, and obtain the ordinary string confusion ratio based on the confusion detection results of all ordinary strings. The ordinary string confusion ratio is the ratio of random strings in ordinary strings.
[0101] Specifically, generally, the length of ordinary strings is relatively long, and it is difficult to have the situation of simple strings. Therefore, for ordinary strings, directly use the random string recognition model to perform confusion detection on them, and then use the ratio of the number of ordinary strings with the confusion detection result of random strings to the total number of ordinary strings as the ordinary string confusion ratio β.
[0102] In an embodiment of the present application, the random string recognition model is a trained CharCNN (Character-level Convolutional Neural Networks) model for identifying whether the input string data is a random string. The structure of the CharCNN model is as Figure 7As shown. When the string to be recognized (in this embodiment, the string to be recognized includes ordinary strings and variable strings that are not simple strings) is input into the CharCNN model, the CharCNN model first performs character-level segmentation processing on the string to be recognized, splitting the string to be recognized into individual characters. Then, according to the pre-constructed character-index mapping table (look-up table), each character is converted into a corresponding one-hot vector (also known as one-hot encoding) through the Embedding layer, obtaining a vector matrix of character embeddings (or called the input matrix). Next, the convolutional layer performs convolution and pooling operations on this input matrix to obtain character features. Then the fully connected layer processes the character features (e.g., using the softmax activation function), obtaining the probability that the string to be recognized is a random string. Finally, this probability is compared with a preset probability threshold to obtain the recognition result of the string to be recognized. For example, if the preset probability threshold is 0.6, then when the probability that the string to be recognized is a random string is greater than 0.6, the output recognition result is a random string; when the probability that the string to be recognized is a random string is less than or equal to 0.6, the output recognition result is a normal string (i.e., a non-random string).
[0103] Step 330: Determine the obfuscation detection result of the source code file according to the variable obfuscation ratio and the ordinary string obfuscation ratio.
[0104] Specifically, when the variable obfuscation ratio and the ordinary string obfuscation ratio of a source code file are determined, it means that the obfuscation detection of all strings in the source code file has been completed. Then, it is possible to determine whether the source code file is an obfuscated code file according to the variable obfuscation ratio and the ordinary string obfuscation ratio.
[0105] In an embodiment of the present application, the steps for determining the obfuscation detection result of the source code file are specifically as follows: when the variable obfuscation ratio is greater than the first preset character threshold and the ordinary string obfuscation ratio is greater than the second preset character threshold, determine that the obfuscation detection result of the source code file is an obfuscated code file; when the variable obfuscation ratio is less than the first preset character threshold or the ordinary string obfuscation ratio is less than the second preset character threshold, determine that the obfuscation detection result of the source code file is an unobfuscated code file.
[0106] Specifically, when both the proportion of variable obfuscation and the proportion of ordinary string obfuscation are greater than their corresponding thresholds, the corresponding source code file is considered an obfuscated code file; otherwise, the corresponding source code file is considered an unobfuscated code file. Exemplarily, it can be expressed as: when α > α' and β > β', the source code file is considered an obfuscated code file; when α ≤ α' or β ≤ β' (that is, at least one of the two conditions α > α' and β > β' is not satisfied), the source code file is considered an unobfuscated code file, where α' represents the first preset character threshold, 0 ≤ α' ≤ 1; and β' represents the second preset character threshold, 0 ≤ β' ≤ 1.
[0107] In an embodiment of the present application, as Figure 4 shown, the application program hardening detection method provided by the present application further includes the generation process of a random string recognition model, specifically including steps 410 to 430, as follows:
[0108] Step 410: Obtain a sample data set, where the sample data set includes sample strings.
[0109] Specifically, based on the characteristics of the code, the strings in the code are composed of pure English characters or pure English characters and other special ASCII (American Standard Code for Information Interchange) characters. Therefore, the sample strings in the sample data set include Chinese pinyin, English words, uppercase English characters, lowercase English characters, and delimiters.
[0110] Chinese pinyin is also a type of pure English character. Incorporating Chinese pinyin into the sample data set is mainly considered that in the Chinese environment, sometimes pure pinyin is used to express some code content. The pinyin corresponding to the Chinese characters in the Chinese character frequency table can be collected as the Chinese pinyin in the sample data set. For example, the pinyin of Chinese characters can be obtained through the Jun Da Chinese character frequency table. This table already contains the pinyin encoding of each Chinese character. Convert all the Chinese pinyin to lowercase characters to form the Chinese character frequency pinyin table.
[0111] English words can be obtained through an English dictionary, and English words with high word frequencies can be preferentially selected. For example, using the COCA (Corpus of Contemporary American English) 60,000-word frequency table, convert all the English words (or high-frequency English words) in the word frequency table to lowercase characters to form the English word frequency table.
[0112] The delimiters are mainly some commonly used special visible characters, such as characters like $, _, -, @, etc. The delimiter table can be formed using ASCII visible characters and null characters.
[0113] Uppercase English characters refer to the 26 uppercase English letters, and lowercase English characters refer to the 26 lowercase English letters.
[0114] The above-mentioned Chinese character frequency and pinyin table, English word frequency table, delimiter table, 26 uppercase English letters, and 26 lowercase English letters constitute the sample data set.
[0115] Step 420: Randomly transform and splice the sample strings in the sample data set to form a training sample set.
[0116] Specifically, the number of samples in the sample data set is limited. To train a neural network model, more sample data is required. In the embodiments of this application, the sample data set is processed by means of data construction to generate a training sample set for training the neural network, without the need to manually collect a large amount of data, reducing the development workload and improving the development efficiency.
[0117] Data construction refers to randomly transforming and splicing sample strings. Random transformation refers to changing the case of sample strings, and splicing processing refers to splicing different types of sample strings to form a new string, which is the sample data in the training sample set.
[0118] In an embodiment of this application, the training sample set includes positive sample data and negative sample data. Positive sample data refers to normal strings, and negative sample data are random strings. The generation of the training sample set includes the generation of positive sample data and the generation of negative sample data.
[0119] In an embodiment of this application, the process of generating positive sample data is specifically as follows: randomly change the case of Chinese pinyin to obtain extended Chinese pinyin; randomly change the case of English words to obtain extended English words; splice and process the extended Chinese pinyin, extended English words, and delimiters to form the positive sample data of the training sample set.
[0120] Specifically, random transformation refers to the random case transformation of Chinese pinyin and English words. Specifically, some characters in Chinese pinyin or English words are changed from uppercase to lowercase, or from lowercase to uppercase. The random case transformation can be: changing Chinese pinyin or English words to all uppercase, changing Chinese pinyin or English words to capitalize the first letter, and changing Chinese pinyin or English words to all lowercase (since the Chinese pinyin and English words obtained in the previous steps are all lowercase, this situation is equivalent to keeping the original sample string). The string after random transformation is called an extended string. For example, for the Chinese pinyin "zhangsan", the extended strings obtained by random transformation are: "ZHANGSAN", "zhangsan", and "Zhangsan". It can be understood that the three random transformation forms described above are determined according to general usage habits. If necessary, more random transformations can be performed. For example, "zhangsan" can be changed to "ZHANGsan", "zhangSAN", "zhangSan", etc.
[0121] After random transformation, the extended Chinese pinyin, extended English words, and delimiters are concatenated to form the positive sample data of the training sample set. When concatenating, preferably, there are delimiters at both ends of each extended Chinese pinyin or extended English word.
[0122] In an embodiment of the present application, the concatenation process of the extended Chinese pinyin, extended English words, and delimiters specifically includes at least one of the following situations:
[0123] Randomly select n extended Chinese pinyin and n + 1 delimiters, and perform an alternating concatenation process on the n extended Chinese pinyin and n + 1 delimiters;
[0124] Randomly select n extended English words and n + 1 delimiters, and perform an alternating concatenation process on the n extended English words and n + 1 delimiters;
[0125] Randomly select m extended Chinese pinyin, n - m extended English words, and n + 1 delimiters, and perform an alternating concatenation process on the n extended Chinese pinyin and extended English words and n + 1 delimiters; m < n, and n, m are natural numbers greater than 0.
[0126] Exemplarily, taking the concatenation of 2 extended Chinese pinyin ("zhangsan", "lisi") and 3 delimiters ("!", "@", "#") as an example, one situation of the alternating concatenation process is: "!zhangsan@lisi#", and it can also be: "@zhangsan!lisi#", "@zhangsan#lisi!", "!lisi@zhangsan#", etc.
[0127] Exemplarily, the generation process of the positive sample data can refer to Figure 5 the flowchart shown. AsFigure 5 As shown, the generation process of positive sample data includes three aspects: the data construction of Chinese pinyin and delimiters, the data construction of English words and delimiters, and the data construction of Chinese pinyin, English words and delimiters.
[0128] As Figure 5 shown, the data construction process of Chinese pinyin and delimiters includes: S511, Chinese character frequency pinyin table (lowercase), S521, pure Chinese pinyin (lowercase), S531, random case transformation (all uppercase, all lowercase, first letter uppercase) to generate n units, S513, all delimiters, S532, randomly select n + 1 delimiters, S540, splice the delimiters and units to generate a normal string.
[0129] As Figure 5 shown, the data construction process of English words and delimiters includes: S512, English word frequency table (lowercase), S523, pure English words (lowercase), S531, random case transformation (all uppercase, all lowercase, first letter uppercase) to generate n units, S513, all delimiters, S532, randomly select n + 1 delimiters, S540, splice the delimiters and units to generate a normal string.
[0130] As Figure 5 shown, the data construction process of English words and delimiters includes: S512, English word frequency table (lowercase), S522, mixed pinyin and words (lowercase), S531, random case transformation (all uppercase, all lowercase, first letter uppercase) to generate n units, S513, all delimiters, S532, randomly select n + 1 delimiters, S540, splice the delimiters and units to generate a normal string.
[0131] Among them, in Figure 5 , the unit in step S531 is equivalent to the extended Chinese pinyin described above, and the normal string in step S540 is equivalent to the positive sample data described above. The specific splicing method can refer to the previous description and will not be elaborated here.
[0132] In an embodiment of the present application, the generation process of negative sample data is specifically: generate n character units according to uppercase English characters and lowercase English characters, and the character units are composed of a random number of English characters selected from uppercase English characters and lowercase English characters; alternately splice the n character units and randomly selected n + 1 delimiters to form the negative sample data of the training sample set.
[0133] Specifically, a character unit is a string composed of a certain number of uppercase English characters or lowercase English characters, and the number is random. That is to say, the length of the character unit is a random length. Exemplarily, if the length of a character unit is set to 4, then the character unit is composed of randomly selected 4 characters from 52 characters including uppercase English characters and lowercase English characters. For each set of random length (or the number of characters in the character unit), a character unit can be generated. Then, by setting n random lengths, n character units can be obtained. Then randomly select n + 1 delimiters, and perform an alternating splicing process on these n character units and n + 1 delimiters to form the negative sample data of the training sample set.
[0134] Exemplarily, the generation process of the negative sample data can refer to Figure 6 , which specifically includes: S611. The English letters a - z and A - Z, that is, 26 uppercase English characters and 26 lowercase English characters. S612. Randomly select m words to form a unit, that is, randomly select m words from 26 uppercase English characters and 26 lowercase English characters to form a character unit, where m is a natural number from 3 to 7. S613. Randomly generate n units, that is, repeat step S612 n times to obtain n character units. S621. All delimiters. S622. Randomly select n + 1 delimiters. S630. Splice the delimiters and the units to generate a random string. That is, alternately splice n character units and n + 1 delimiters to generate a random string, and this random string is the negative sample data in the training sample set.
[0135] Continue to refer to Figure 4 , step 430. Train a neural network model based on the training sample set to obtain a random string recognition model.
[0136] Specifically, train the neural network model based on the training sample set so that the neural network model can accurately distinguish between random strings and non - random strings (i.e., normal strings). The trained neural network model is the random string recognition model.
[0137] In an embodiment of the present application, since the sample data in the training sample set is randomly transformed and spliced, the string lengths of the sample data (abbreviated as the lengths of the sample data) may be different, while the neural network model requires the lengths of the input data to be consistent. Therefore, when training the neural network model based on the training sample set, it is also necessary to pre - process the training sample set to form the input data set of the neural network model; then train the neural network model based on the input data set to obtain a random string recognition model.
[0138] The preprocessing of the training sample set specifically includes: determining the maximum length of the sample data in the training sample set, where the length of the sample data refers to the number of characters contained in the sample data; splicing the sample data with a length less than the maximum length with a preset character to form input data with a length reaching the maximum length. Specifically, the maximum length of the sample data in the training sample set is used as the reference length of the input data of the neural network model, and then the sample data with a length less than the maximum length is spliced with a preset character so that the length of the sample data reaches the maximum length. In this way, it can be ensured that the lengths of all data in the input data set are the same. When splicing the sample data with a smaller length with the preset character, according to the difference between the length of the sample data and the maximum length, the corresponding number of preset characters are spliced. For example, if the length of a sample data is 5 and the maximum length is 7, then 2 preset characters need to be spliced at the end of the sample data.
[0139] In one embodiment of the present application, the neural network model adopts a CharCNN (Character-level Convolutional Neural Networks) model. The processing process of the CharCNN model for the input data is as Figure 7 shown. As Figure 7 shown, the CharCNN model includes an Embedding (embedding layer), a convolutional layer, and a fully connected layer. The input data (i.e., the input text) is "Hello", the length of this data is 5 characters, while the length of the input data of the CharCNN model is 7 characters, so it is preprocessed, and two preset characters " <pad>". After preprocessing, the input data undergoes char embedding operation on characters through the embedding layer, that is, each character is converted into a corresponding one-hot vector to obtain the input matrix. Then, operations are performed on the input matrix through network layers such as convolutional layers and fully connected layers to obtain the output result. Among them, the processing operations of the convolutional layer include performing convolution and pooling operations on the input matrix to obtain character features. Then, the fully connected layer calculates the character features (for example, using the softmax activation function for processing) to obtain the output data. After obtaining the output data, the output data is compared and calculated with the training target to obtain the loss function, and the model parameters are updated based on the loss function. Then, the next input data is processed according to the model with updated parameters, and so on in a loop until the loss function meets the conditions or the number of loops reaches the set number, then the trained CharCNN model, that is, the random string recognition model, is obtained.
[0140] In an embodiment of the present application, during the training process of the CharCNN model, a part of the data is extracted from the training sample set as the test data set. After obtaining the trained CharCNN model, the trained CharCNN model is tested through the test data set, that is, the accuracy of the trained CharCNN model for string recognition is tested. When the accuracy reaches the preset value, it indicates that the trained CharCNN model can be used as a random string recognition model; when the accuracy does not reach the preset value, it indicates that the string recognition accuracy of the trained CharCNN model is relatively low and needs to be improved. For example, the model can be retrained by adjusting model hyperparameters, increasing the number of training times, etc.
[0141] In an embodiment of the present application, the neural network model can also use other suitable natural language processing models or machine learning models. Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language used by people in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph, and other technologies.
[0142] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0143] Continue to refer to Figure 2 , step 240: Determine whether the application is fortified according to the obfuscation detection results of all source code files.
[0144] Specifically, according to the obfuscation detection results of all source code files, when the number of source code files with obfuscated code files reaches a preset threshold, it is considered that the application has been fortified; when the number of source code files with obfuscated code files is less than the preset threshold, it is considered that the application has not been fortified.
[0145] In an embodiment of the present application, it is also possible to determine whether the application is fortified based on the number of obfuscated code files (that is, the number of source code files with obfuscated code files in the obfuscation detection results) and the number of non-obfuscated code files (the number of source code files with non-obfuscated code files in the obfuscation detection results). When the number of obfuscated code files is greater than the number of non-obfuscated code files, it is considered that the application has been fortified; conversely, when the number of obfuscated code files is less than or equal to the number of non-obfuscated code files, it is considered that the application has not been fortified.
[0146] In the technical solution provided by the embodiment of the present application, by obtaining multiple source code files of the application, then parsing each source code file respectively, and performing obfuscation detection on each source code file based on the parsing results, and finally determining whether the application is fortified according to the obfuscation detection results of all source code files; dividing the application into multiple source code files for obfuscation detection, that is, refining the detection granularity of the fortification detection and improving the accuracy of the fortification detection.
[0147] In an embodiment of the present application, after parsing the source code file in step 220, interface call data of the source code file can also be obtained. Then, the application fortification detection method provided by the present application further includes: performing logical detection on the source code file according to the interface call data of the source code file to obtain a logical detection result of the source code file, and the logical detection result is a logical hidden code file or a logical non-hidden code file.
[0148] Specifically, the interface call data of the source code file refers to the code instructions in the source code that represent the call to the interface (Application Program Interface, API). For example, wx.getUserInfo represents the call to the user information API.
[0149] Logical detection of the source code file is to detect whether the interface call method in the source code file is hidden. Hiding the interface call means that the code does not contain the official api call code, but transforms the interface name into other forms (such as setting it as a variable) and then makes the call. This call method is called implicit interface call. In contrast to the implicit interface call, making an interface call through the official api call code is called explicit interface call. For example, the interface call code includes wx.getUserInfo.
[0150] The logically hidden code file means that the interface call in the source code file is hidden, that is, the implicit interface call method is mostly used. The logically unhidden code file means that the interface call in the source code file is not hidden, that is, the explicit interface call method is mostly used. Therefore, by judging the number of implicit and explicit interface calls in the interface call data, the logical detection result of the source code file can be determined.
[0151] In an embodiment of the present application, the process of logically detecting the source code file is specifically as follows: determining the proportion of implicit interface calls in the source code file according to the explicit interface call data and implicit interface call data of the source code file; when the proportion of implicit interface calls is less than the preset interface threshold, determining that the logical detection result of the source code file is a logically unhidden code file; when the proportion of implicit interface calls is greater than the preset interface threshold, determining that the logical detection result of the source code file is a logically hidden code file.
[0152] Specifically, count the number of explicit interface calls c according to the explicit interface call data 1 , count the number of implicit interface calls c according to the implicit interface call data 2 , and the calculation method of the proportion δ of implicit interface calls is as follows:
[0153]
[0154] If δ’ represents the preset interface threshold, then when δ ≤ δ’, determine that the logical detection result of the source code file is a logically unhidden code file; when δ > δ’, determine that the logical detection result of the source code file is a logically hidden code file.
[0155] In an embodiment of the present application, the obfuscation detection and logic detection of the source code file can be performed synchronously. When both the obfuscation detection and the logic detection are performed on the source code file, step 240 is: determining whether the application is fortified according to the obfuscation detection results of all source code files and the logic detection results of all source code files.
[0156] Specifically, according to the obfuscation detection results of all source code files, it can be determined whether the application has code obfuscation. According to the logic detection results of all source code files, it can be determined whether the application has logic hiding. When it is determined that the application has both code obfuscation and logic hiding, it can be considered that the application is fortified. It can be understood that when the application meets one of the conditions of code obfuscation and logic hiding, it can also be considered that the application is fortified, but the degree of fortification is weaker.
[0157] In an embodiment of the present application, the process of determining whether the application is fortified specifically includes: determining the proportion of obfuscated code files according to the number of source code files that are obfuscated code files among all source code files and the total number of source code files; determining the proportion of logic-hidden code files according to the number of source code files that are logic-hidden code files among all source code files and the total number of source code files; when the proportion of obfuscated code files is greater than the first preset file threshold and the proportion of logic-hidden code files is greater than the second preset file threshold, determining that the application has been fortified; when the proportion of obfuscated code files is less than the first preset file threshold or the proportion of logic-hidden code files is less than the second preset file threshold, determining that the application has not been fortified.
[0158] Specifically, the proportion of obfuscated code files in the source code files is the proportion of obfuscated code files γ, and the proportion of logic-hidden code files in the source code files is the proportion of logic-hidden code files ξ. γ' represents the first preset file threshold, and ξ' represents the second preset file threshold. Then, when γ > γ' and ξ > ξ', it is considered that the application has been fortified; when γ ≤ γ' or ξ ≤ ξ', it is considered that the application has not been fortified.
[0159] In the technical solution provided by the embodiment of the present application, it is realized to judge whether the application code project is fortified from two aspects of obfuscation detection and logic detection, and the judgment of code fortification is more comprehensive and the judgment result is more accurate.
[0160] In an embodiment of the present application, when it is determined that the application has not been fortified, corresponding prompt information can be generated so that developers can perform fortification operations on the application to improve the security of the application.
[0161] Exemplarily, Figures 8 - 10 Taking an application as a small program as an example to illustrate the specific implementation process of the technical solution of this application. A small program is a hosted program that runs in a host program. On the user side, the small program does not need to be downloaded and installed, and the user can use the small program through the host program. When performing reinforcement detection on the hosted program, the hosted program to be detected can be determined first in the host program. For example, the hosted program that has been created in the host program can be used as the hosted program to be detected, and the hosted program to be detected can be determined through the APPID (program identifier), program name, etc. The APPID is the unique identifier assigned by the host program to the hosted program. When the host program opens the hosted program, it usually needs to download the code package of the hosted program to the local. Therefore, after determining the hosted program to be detected, the code package of the hosted program to be detected can be obtained through the host program, and the content of this code package can be used as the source code file of the hosted program.
[0162] Figure 8 The flowchart showing whether the code of a small program is obfuscated provided by an embodiment of this application is shown, that is, the flowchart of small program obfuscation detection. As Figure 8 shown, the process of determining whether the code of a small program is obfuscated includes steps S810 to S850, which are specifically as follows:
[0163] S810. Obtain the small program code project, which includes the executable code file of the small program.
[0164] S820. Decompile and obtain all js files. The source code of the small program is written in the Javascript language, and the source code file obtained by decompiling the executable code file is also called a Javascript file, simply referred to as a js file.
[0165] S830, js file obfuscation judgment module. In this step, obfuscation detection is performed on each js file. The process of obfuscation detection includes steps S831 to S834, specifically: S831, perform AST syntax tree parsing on the Javascript file to obtain all variables and strings. Among them, variables are equivalent to the variable strings described above, and strings are equivalent to the ordinary strings described above. S832, perform obfuscation detection on all variables through a simple character recognition algorithm and a random string recognition model to obtain the variable obfuscation ratio. Specifically, first perform simple character recognition on the variables. When the variable is a simple character, it is considered that the variable is a random string; when the variable is not a simple character, judge whether the variable is a random string through the random string recognition model. Finally, take the ratio of the number of variables determined to be random strings to the total number of variables as the variable obfuscation ratio α. The specific content of step S832 can refer to the relevant description in step 230 above and will not be elaborated here. S833, perform obfuscation detection on all strings through the random string recognition model to obtain the string obfuscation encryption ratio. The string obfuscation encryption ratio is the ordinary string obfuscation ratio β described above. For specific reference, refer to the relevant description in step 230 above and will not be elaborated here. S834, judge whether the variable obfuscation ratio and the string obfuscation encryption ratio exceed the corresponding thresholds. If the variable obfuscation ratio is greater than the first preset character threshold and the string obfuscation encryption ratio is greater than the second preset character threshold, it is considered that the js file is obfuscated; otherwise, it is considered that the js file is not obfuscated.
[0166] S840, calculate the file obfuscation ratio. The file obfuscation ratio is the obfuscated code file ratio γ described above. Divide the number of js files determined to be obfuscated by the total number of js files to obtain the file obfuscation ratio.
[0167] S850, judge whether the file obfuscation ratio exceeds the corresponding threshold. In this embodiment, the thresholds include the first threshold γ 1 and the second threshold γ 2 , when the file obfuscation ratio γ ≤ γ 1 , it is considered that the applet is not obfuscated (i.e., the project is not obfuscated); when the file obfuscation ratio γ 1 <γ ≤ γ 2 , it is considered that the applet is weakly obfuscated (i.e., the project is weakly obfuscated); when the file obfuscation ratio γ > γ 2 , it is considered that the applet is strongly obfuscated (i.e., the project is strongly obfuscated).
[0168] Figure 9 shows the flowchart for determining whether an applet performs logical hiding provided by an embodiment of the present application, that is, the flowchart for applet logic detection. As Figure 9 shown, the process of determining whether an applet performs logical hiding includes steps S910 to S950, specifically as follows:
[0169] S910. Obtain the mini-program code project, which includes the executable code files of the mini-program.
[0170] S920. Decompile and obtain all js files. The source code of the mini-program is written in the Javascript language. The source code file obtained by decompiling the executable code file is also called a Javascript file, simply referred to as a js file.
[0171] S930. js file logic hiding judgment module. In this step, logical detection is performed on each js file. The process of logical detection includes steps S931 to S934. Specifically: S931. Parse the AST syntax tree of the Javascript file to obtain the function call chain, and then obtain all wx api calls. Among them, all wx api calls are the interface call data described above, including display interface call data and implicit interface call data. S932. Count the display calls and implicit calls of wx api. Display calls refer to the number of display interface calls, and implicit calls refer to the number of implicit interface calls. S933. Calculate the proportion of implicit calls of wx api. The proportion of implicit calls is equivalent to the proportion of implicit interface calls δ described above. The specific calculation method refers to the relevant description in step 240 above and will not be elaborated here. S934. Judge whether the proportion of implicit calls exceeds the corresponding threshold. If the proportion of implicit calls is greater than the preset interface threshold, it is considered that the code logic is hidden; if the proportion of implicit calls is less than or equal to the preset interface threshold, it is considered that the code logic is not hidden.
[0172] S940. Calculate the proportion of file logic hiding. The proportion of file logic hiding is the proportion of logic hidden code files ξ described above. Divide the number of js files determined to have hidden code logic by the total number of js files to obtain the proportion of file logic hiding.
[0173] S950. Judge whether the proportion of file logic hiding exceeds the corresponding threshold. When the proportion of file logic hiding exceeds the corresponding threshold, it is considered that the project logic is hidden; when the proportion of file logic hiding does not exceed the corresponding threshold, it is considered that the project logic is not hidden.
[0174] Figure 10 shows a flowchart of the mini-program reinforcement detection method provided by an embodiment of the present application. As Figure 10 shown, the mini-program reinforcement detection method is specifically as follows:
[0175] S1010. Obtain the executable code file of the mini-program and decompile it to obtain a js file. The source code of the mini-program is written in the JavaScript language, and the source code file obtained by decompiling the executable code file is also called a JavaScript file, simply referred to as a js file.
[0176] S1020. Parse the AST syntax tree and obtain variable names and strings. Parse each js file through an AST syntax tree parsing tool to obtain the variable strings and ordinary strings corresponding to the js, where the variable strings refer to Figure 10 the variable names therein, and the ordinary strings refer to Figure 10 the strings therein. After parsing, the interface call data corresponding to the js file can also be obtained. After parsing, steps S1030 and S1040 can be performed simultaneously.
[0177] S1030. Perform obfuscation detection on the mini-program according to the variable names and strings of the js file, including steps S1031 to S1035, specifically: S1031. Obtain a sample data set, which includes: a Chinese character frequency pinyin table (lowercase), an English word frequency table (lowercase), and a delimiter table, and also includes 26 uppercase English letters and 26 lowercase English letters. S1032. Perform random transformation and splicing processing on the sample strings in the sample data set to form a training sample set, which includes normal strings and random strings. The normal strings are the positive sample data of the training sample set, and the random strings are the negative sample data of the training sample set. The generation process of the training sample set can refer to the relevant description in step 420 above and will not be elaborated here. S1033. Train the CharCNN model. When training the CharCNN model, preprocessing operations also need to be performed on the sample data in the training sample set. This step can specifically refer to the relevant description in step 430 above and will not be elaborated here. S1034. Use the CharCNN model to judge the randomness of variable names and strings. This step is to perform obfuscation detection on the variable names and strings of the js file parsed in step S1020. Specifically, it can refer to the relevant description in step 230 above and will not be elaborated here. S1035. Judge whether the file obfuscation ratio exceeds the corresponding threshold. The file obfuscation ratio is the ratio of the obfuscated js files in the js file, that is, the obfuscated code file ratio mentioned above. In this embodiment, the threshold includes the first threshold γ 1 and the second threshold γ 2 . When the file obfuscation ratio γ ≤ γ 1 , it is considered that the mini-program is not obfuscated (it can be simply recorded as the mini-program is not obfuscated); when the file obfuscation ratio γ 1 <γ ≤ γ 2 , it is considered that the mini-program is weakly obfuscated (it can be simply recorded as the mini-program is weakly obfuscated); when the file obfuscation ratio γ > γ 2 When it is, it is considered that the applet is strongly obfuscated (which can be simply recorded as the applet is strongly obfuscated).
[0178] S1040. Detect the applet logic according to the interface call data of the js file. Specifically: Detect the proportion of implicit calls of the wx official api in the JavaScript code. When the proportion of implicit calls exceeds the corresponding threshold, it is considered that the applet logic is not hidden; when the proportion of implicit calls does not exceed the corresponding threshold, it is considered that the applet logic is hidden. The proportion of implicit calls is equivalent to the proportion of the logic hidden code files described above. For specific reference, please refer to the relevant description in step 240 above, which will not be elaborated here.
[0179] S1050. If the result of the obfuscation detection is weak obfuscation and strong obfuscation, and the result of the logic detection is that the logic is hidden, it is determined that the applet is fortified. In other cases, it is considered that the applet is not fortified.
[0180] It should be noted that although the steps of the method in this application are described in a specific order in the drawings, this does not require or imply that these steps must be executed in this specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0181] The following introduces the device embodiments of this application, which can be used to execute the application hardening detection method in the above embodiments of this application. Figure 11 Schematically shows the structural block diagram of the application hardening detection device provided by the embodiments of this application. As Figure 11 shown, the application hardening detection device provided by the embodiments of this application includes:
[0182] A source code acquisition module 1110, configured to acquire multiple source code files of an application;
[0183] A code parsing module 1120, configured to parse each of the source code files to obtain multiple strings corresponding to the source code files;
[0184] An obfuscation detection module 1130, configured to perform obfuscation detection by determining each string corresponding to the source code file, and obtain an obfuscation detection result of the source code file, where the obfuscation detection result is an obfuscated code file or an unobfuscated code file;
[0185] A code hardening determination module 1140, configured to determine whether the application is hardened according to the obfuscation detection results of all source code files.
[0186] In an embodiment of this application, the string includes a variable string and a normal string; the obfuscation detection module 1130 includes:
[0187] A variable string detection unit, configured to perform obfuscation detection on the variable string through a simple character recognition algorithm and a random string recognition model, and obtain a variable obfuscation ratio according to the obfuscation detection results of all variable strings, where the variable obfuscation ratio is the ratio of the random string in the variable string;
[0188] A common string detection unit, configured to perform obfuscation detection on the common string through the random string recognition model, and obtain a common string obfuscation ratio according to the obfuscation detection results of all common strings, where the common string obfuscation ratio is the ratio of the random string in the common string;
[0189] An obfuscation result determination unit, configured to determine the obfuscation detection result of the source code file according to the variable obfuscation ratio and the common string obfuscation ratio.
[0190] In an embodiment of the present application, the variable string detection unit is specifically configured to:
[0191] Determine whether the variable string is a simple string through the simple character recognition algorithm; if the variable string is a simple string, determine that the variable string is a random string; if the variable string is not a simple string, determine whether the variable string is a random string through the random string recognition model; determine the variable obfuscation ratio according to the ratio of the number of variable strings determined to be random strings to the total number of variable strings.
[0192] In an embodiment of the present application, the obfuscation result determination unit is specifically configured to:
[0193] When the variable obfuscation ratio is greater than a first preset character threshold and the common string obfuscation ratio is greater than a second preset character threshold, determine that the obfuscation detection result of the source code file is an obfuscated code file; when the variable obfuscation ratio is less than the first preset character threshold or the common string obfuscation ratio is less than the second preset character threshold, determine that the obfuscation detection result of the source code file is an unobfuscated code file.
[0194] In an embodiment of the present application, the device further includes:
[0195] A sample set acquisition module, configured to acquire a sample data set, where the sample data set includes sample strings;
[0196] A training set generation module, configured to perform random transformation and splicing processing on the sample strings in the sample data set to form a training sample set;
[0197] A model training module, configured to train a neural network model based on the training sample set to obtain the random string recognition model.
[0198] In an embodiment of the present application, the sample string includes Chinese pinyin, English words, and delimiters, and the training sample set includes positive sample data; the training set generation module includes:
[0199] A first random transformation unit, configured to perform random case transformation on the Chinese pinyin to obtain extended Chinese pinyin; perform random case transformation on the English words to obtain extended English words;
[0200] A first splicing unit, configured to splice the extended Chinese pinyin, the extended English words, and the delimiters to form the positive sample data of the training sample set.
[0201] In an embodiment of the present application, the first random transformation unit is configured to perform at least one of the following operations:
[0202] Randomly select n extended Chinese pinyins and n + 1 delimiters, and alternately splice the n extended Chinese pinyins and n + 1 delimiters;
[0203] Randomly select n extended English words and n + 1 delimiters, and alternately splice the n extended English words and n + 1 delimiters;
[0204] Randomly select m extended Chinese pinyins, n - m extended English words, and n + 1 delimiters, and alternately splice the n extended Chinese pinyins, the extended English words, and n + 1 delimiters; m < n, and n and m are natural numbers greater than 0.
[0205] In an embodiment of the present application, the sample string includes capital English characters, lowercase English characters, and delimiters, and the training sample set includes negative sample data; the training set generation module includes:
[0206] A second random transformation unit, configured to generate n character units according to the capital English characters and the lowercase English characters, and one character unit is composed of a random number of English characters selected from the capital English characters and the lowercase English characters;
[0207] A second splicing unit, configured to alternately splice the n character units and randomly selected n + 1 delimiters to form the negative sample data of the training sample set.
[0208] In an embodiment of the present application, the model training module includes:
[0209] A data preprocessing unit for preprocessing the training sample set to form an input data set for the neural network model; the preprocessing includes: determining the maximum length of the sample data in the training sample set, where the length of the sample data refers to the number of characters contained in the sample data; concatenating the sample data with a length less than the maximum length with a preset character to form input data with a length reaching the maximum length.
[0210] A model training unit for training the neural network model based on the input data set to obtain the random string recognition model.
[0211] In an embodiment of the present application, after parsing the source code file, interface call data in the source code file is also obtained; the apparatus further includes:
[0212] A logic detection module for performing logic detection on the source code file according to the interface call data of the source code file to obtain a logic detection result of the source code file, where the logic detection result includes a logic hidden code file and a logic unhidden code file.
[0213] Correspondingly, the code hardening determination module 1140 is used for:
[0214] Determining whether to harden the application according to the obfuscation detection results of all source code files and the logic detection results of all source code files.
[0215] In an embodiment of the present application, the interface call data includes explicit interface call data and implicit interface call data; the logic detection module is specifically used for:
[0216] Determining the proportion of implicit interface calls in the source code file according to the explicit interface call data and implicit interface call data of the source code file; when the proportion of implicit interface calls is less than a preset interface threshold, determining the logic detection result of the source code file as a logic unhidden code file; when the proportion of implicit interface calls is greater than the preset interface threshold, determining the logic detection result of the source code file as a logic hidden code file.
[0217] In an embodiment of the present application, the code hardening determination module 1140 is specifically used for:
[0218] Determine the proportion of obfuscated code files based on the number of source code files of the obfuscated code files and the total number of source code files in the obfuscation detection result; determine the proportion of logic-hidden code files based on the number of source code files of the logic-hidden code files and the total number of source code files in the logic detection result; when the proportion of the obfuscated code files is greater than the first preset file threshold and the proportion of the logic-hidden code files is greater than the second preset file threshold, determine that the application has been fortified; when the proportion of the obfuscated code files is less than the first preset file threshold or the proportion of the logic-hidden code files is less than the second preset file threshold, determine that the application has not been fortified.
[0219] The specific details of the application fortification detection device provided in each embodiment of the present application have been described in detail in the corresponding method embodiments, and will not be repeated here.
[0220] Figure 12 Schematically shows a block diagram of a computer system of an electronic device for implementing the embodiments of the present application.
[0221] It should be noted that Figure 12 The computer system 1200 of the shown electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0222] As Figure 12 shown, the computer system 1200 includes a central processing unit 1201 (Central Processing Unit, CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 1202 (Read-Only Memory, ROM) or the program loaded from the storage section 1208 into the random access memory 1203 (Random Access Memory, RAM). In the random access memory 1203, various programs and data required for system operation are also stored. The central processing unit 1201, the read-only memory 1202, and the random access memory 1203 are connected to each other through a bus 1204. The input / output interface 1205 (Input / Output interface, that is, I / O interface) is also connected to the bus 1204.
[0223] The following components are connected to the input / output interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a local area network card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the input / output interface 1205 as needed. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as needed so that a computer program read from it can be installed into the storage section 1208 as needed.
[0224] Specifically, according to an embodiment of the present application, the processes described in each method flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by the central processing unit 1201, various functions defined in the system of the present application are executed.
[0225] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or process a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be processed by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0226] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0227] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0228] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0229] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known knowledge or conventional technical means in the technical field not disclosed in the present application.
[0230] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.< / pad>
Claims
1. An application program reinforcement detection method, characterized in that, it includes: Obtain multiple source code files of the application program; Parse each of the source code files to obtain multiple strings and interface call data corresponding to the source code files; Obtain a sample data set, where the sample data set includes sample strings; Perform random transformation and splicing processing on the sample strings in the sample data set to form a training sample set; Train a neural network model based on the training sample set to obtain a random string recognition model; Determine whether each string corresponding to the source code file is a random string through the random string recognition model to perform confusion detection on the source code file, and obtain the confusion detection result of the source code file, where the confusion detection result is a confused code file or an unconfused code file; Perform logical detection on the source code file according to the interface call data of the source code file to obtain the logical detection result of the source code file, where the logical detection result includes a logically hidden code file and a logically unhidden code file; Determine whether the application program is reinforced according to the confusion detection results of all source code files and the logical detection results of all source code files; Among them, performing random transformation and splicing processing on the sample strings in the sample data set to form a training sample set includes: Perform random case transformation on the Chinese pinyin included in the sample string to obtain extended Chinese pinyin; perform random case transformation on the English words included in the sample string to obtain extended English words; Perform splicing processing on the extended Chinese pinyin, the extended English words, and the delimiters included in the sample string to form the positive sample data of the training sample set.
2. The application program reinforcement detection method according to claim 1, characterized in that, the string includes a variable string and a common string; Performing confusion detection on the source code file by determining whether each string corresponding to the source code file is a random string through the random string recognition model includes: Performing confusion detection on the variable string through a simple character recognition algorithm and the random string recognition model, and obtaining a variable confusion ratio according to the confusion detection results of all variable strings, where the variable confusion ratio is the ratio of random strings in the variable string; Performing confusion detection on the common string through the random string recognition model, and obtaining a common string confusion ratio according to the confusion detection results of all common strings, where the common string confusion ratio is the ratio of random strings in the common string; Determine the confusion detection result of the source code file according to the variable confusion ratio and the common string confusion ratio.
3. The application program reinforcement detection method according to claim 2, characterized in that, Performing confusion detection on the variable string through a simple character recognition algorithm and the random string recognition model, and obtaining a variable confusion ratio according to the confusion detection results of all variable strings, includes: Determine whether the variable string is a simple string through the simple character recognition algorithm; If the variable string is a simple string, determine that the variable string is a random string; If the variable string is not a simple string, determine whether the variable string is a random string through the random string recognition model; Determine the variable obfuscation ratio according to the ratio of the number of variable strings determined to be random strings to the total number of variable strings.
4. The application program reinforcement detection method according to claim 2, characterized in that determine the obfuscation detection result of the source code file according to the variable obfuscation ratio and the ordinary string obfuscation ratio, including: When the variable obfuscation ratio is greater than the first preset character threshold and the ordinary string obfuscation ratio is greater than the second preset character threshold, determine that the obfuscation detection result of the source code file is an obfuscated code file; When the variable obfuscation ratio is less than the first preset character threshold or the ordinary string obfuscation ratio is less than the second preset character threshold, determine that the obfuscation detection result of the source code file is an unobfuscated code file.
5. The application program reinforcement detection method according to claim 1, characterized in that perform splicing processing on the extended Chinese pinyin, the extended English words and the delimiters included in the sample string, including at least one of the following situations: Randomly select n extended Chinese pinyins and n + 1 delimiters, and perform alternating splicing processing on the n extended Chinese pinyins and n + 1 delimiters; Randomly select n extended English words and n + 1 delimiters, and perform alternating splicing processing on the n extended English words and n + 1 delimiters; Randomly select m extended Chinese pinyins, n - m extended English words and n + 1 delimiters, and perform alternating splicing processing on the n extended Chinese pinyins and extended English words and n + 1 delimiters; m < n, and n and m are natural numbers greater than 0.
6. The application program reinforcement detection method according to claim 1, characterized in that perform random transformation and splicing processing on the sample strings in the sample dataset to form a training sample set, further including: Generate n character units according to the capital English characters included in the sample string and the lowercase English characters included in the sample string. One character unit is composed of a random number of English characters selected from the capital English characters and the lowercase English characters; Perform alternating splicing processing on the n character units and n + 1 delimiters randomly selected from the sample string to form the negative sample data of the training sample set.
7. The application program reinforcement detection method according to claim 1, characterized in that Train a neural network model based on the training sample set to obtain a random string recognition model, including: Preprocess the training sample set to form the input dataset of the neural network model; the preprocessing includes: determining the maximum length of the sample data in the training sample set, where the length of the sample data refers to the number of characters included in the sample data; splicing the sample data with a length less than the maximum length with a preset character to form input data with a length reaching the maximum length; Train the neural network model based on the input data set to obtain the random string recognition model.
8. The application program hardening detection method according to claim 1, wherein, the interface call data includes explicit interface call data and implicit interface call data; Performing a logic detection on the source code file according to the interface call data of the source code file to obtain a logic detection result of the source code file, including: Determining the proportion of implicit interface calls of the source code file according to the explicit interface call data and the implicit interface call data of the source code file; When the proportion of implicit interface calls is less than a preset interface threshold, determining that the logic detection result of the source code file is a logic non-hidden code file; When the proportion of implicit interface calls is greater than a preset interface threshold, determining that the logic detection result of the source code file is a logic hidden code file.
9. The application program hardening detection method according to claim 1, wherein, Determining whether the application program is hardened according to the obfuscation detection results of all source code files and the logic detection results of all source code files, including: Determining the proportion of obfuscated code files according to the number of source code files with obfuscation detection results being obfuscated code files and the total number of source code files; Determining the proportion of logic hidden code files according to the number of source code files with logic detection results being logic hidden code files and the total number of source code files; When the proportion of obfuscated code files is greater than a first preset file threshold and the proportion of logic hidden code files is greater than a second preset file threshold, determining that the application program has been hardened; When the proportion of obfuscated code files is less than the first preset file threshold, or the proportion of logic hidden code files is less than the second preset file threshold, determining that the application program is not hardened.
10. An application program hardening detection device, wherein, comprising: A source code acquisition module, configured to acquire a plurality of source code files of an application program; A code parsing module, configured to parse each of the source code files to obtain a plurality of strings and interface call data corresponding to the source code file; A sample set acquisition module, configured to acquire a sample data set, where the sample data set includes sample strings; A training set generation module, configured to perform random transformation and splicing processing on the sample strings in the sample data set to form a training sample set; A model training module, configured to train a neural network model based on the training sample set to obtain a random string recognition model; An obfuscation detection module, configured to perform an obfuscation detection on the source code file by determining whether each string corresponding to the source code file is a random string through the random string recognition model, to obtain an obfuscation detection result of the source code file, where the obfuscation detection result is an obfuscated code file or an unobfuscated code file; A logic detection module, configured to perform a logic detection on the source code file according to the interface call data of the source code file to obtain a logic detection result of the source code file, where the logic detection result includes a logic hidden code file and a logic non-hidden code file; A code hardening determination module, configured to determine whether the application is hardened according to the obfuscation detection results of all source code files and the logic detection results of all source code files; Wherein, the training set generation module is specifically configured to: Randomly change the case of the Chinese pinyin included in the sample string to obtain extended Chinese pinyin; randomly change the case of the English words included in the sample string to obtain extended English words; Perform splicing processing on the extended Chinese pinyin, the extended English words, and the delimiters included in the sample string to form the positive sample data of the training sample set.
11. A computer-readable medium, on which a computer program is stored, Characterized in that, When the computer program is executed by a processor, it implements the application hardening detection method according to any one of claims 1 to 9.
12. An electronic device, Characterized in that, Comprising: A processor; And A memory, configured to store executable instructions of the processor; Wherein, when the processor executes the executable instructions, the electronic device executes the application hardening detection method according to any one of claims 1 to 9.
13. A computer program product, Characterized in that, The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; The processor of the computer device reads and executes the computer instructions from the computer-readable storage medium, so that the computer device executes the application hardening detection method according to any one of claims 1 to 9.
Citation Information
Patent Citations
API name and immediate value-based heuristic sample detection method and system
CN105740706A
Script heuristic detection method and system based on variable name confusion degree
CN106650449A