Method suitable for APP analysis SaaS multi-engine analysis
Through the SaaS-based multi-engine APP analysis method, comprehensive analysis and formulation of judgment standards are solved, and the problems of inaccurate analysis results and large resource overhead in the existing APP analysis methods are achieved, and high accuracy and reliability analysis results are achieved, which improves the detection efficiency.
Patent Information
- Application Number
- CN202411752537.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-05-30
AI Technical Summary
The existing APP analysis methods have problems such as inaccurate analysis results, large resource overhead, and equipment environment impact, which makes it difficult to distinguish the correctness of the investigation direction and the analysis results.
The SaaS-based multi-engine analysis method is adopted to comprehensively analyze and formulate analysis and judgment standards through multi-engine comprehensive analysis to ensure that the analysis capabilities keep pace with the times, reduce the adverse effects of the network and device environment, and improve the accuracy of the analysis results through human intervention and self-iteration improvement.
It realizes high accuracy and reliability of analysis results, reduces false positives and omissions, improves investigation efficiency, and lowers the threshold for the analysis process.
Smart Images

Figure CN120068063A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication forensics, and in particular, mainly relates to a method suitable for SaaS - based multi - engine parsing of APP analysis. Background Art
[0002] In recent years, with the rapid development of the mobile Internet, a large number of Internet illegal acts have emerged. At present, network tracing basically collects fraud - related APPs first and analyzes them. In order to quickly crack down on such illegal APPs, it is necessary to comprehensively detect the background behaviors of mobile APPs and find out key information such as background servers and developer identities. The existing APP parsing methods are mainly divided into static parsing and dynamic parsing. Among them, static parsing refers to the process of obtaining information and characteristics about an application program by statically analyzing and checking the code of the application program without actually running the application program. It includes aspects such as syntax analysis, symbol table and type checking, control flow analysis, data flow analysis, detecting code defects and security vulnerabilities, generating reports and warnings, etc. Dynamic parsing refers to the process of obtaining information and characteristics about an application program by monitoring and analyzing the runtime behavior of the application program when it is actually running. It includes aspects such as runtime data collection, runtime behavior analysis, security vulnerability detection, performance analysis, debugging and troubleshooting, etc.
[0003] At present, the APP parsing capabilities on the market vary widely and there is no unified standard. Some perform decompilation analysis, and some perform unpacking analysis. Different parsing programs may lead to very different parsing results. For example, static parsing may produce false positives and omissions, and it is necessary to combine other testing methods or manual code review to comprehensively evaluate the quality and security of the application program. Dynamic parsing will have certain performance and resource overhead on the execution of the application program, and is restricted by the installation environment. Different operating systems may cause installation failures or running failures, and it is also impossible to detect problems related to static behaviors.
[0004] At present, the parsing capabilities of most products are integrated on the device. The network connection status of a single device may affect the packet capture result. Moreover, fraud - related APPs are developing rapidly with fast technological iteration, and the parsing capabilities of devices are basically constant and need to be updated regularly. Otherwise, there will be situations of missed parsing or inability to parse. And incorrect parsing results will seriously affect the investigation direction. APP parsing has a certain threshold, requiring analysts to have a certain degree of professionalism. Therefore, it is often impossible to distinguish the correctness of the parsing results. Summary of the Invention
[0005] Aiming at the above - mentioned deficiencies, the present invention proposes a SaaS - based multi - engine parsing method for APP analysis, which uses multiple engines for comprehensive analysis and formulates analysis and determination criteria, ensuring that the parsing capabilities keep up with the times and are continuously updated.
[0006] The method of the present invention also greatly reduces the adverse effects caused by different network scenarios and different device environments, integrates multiple engines and parsing methods, and formulates parsing rules, avoiding the incorrect parsing results caused by the failure of single-engine parsing to the greatest extent. Coupled with human parsing intervention and continuous self-iterative collection and improvement, it also ensures the loose coupling of parsing and analysis results and ensures the highest degree of accuracy of parsing results, specifically including:
[0007] Use SaaS for distributed system deployment, introduce a parsing engine to obtain parsing results. Among them, the engines include: basic application parsing, decompilation analysis, unpacking analysis, virus analysis, website analysis, and script analysis. Calculate the weight factors of each engine using the parsing results, specifically including: collect the APP parsing results that have been verified in actual combat as samples and import them into the corresponding engine for parsing, compare the engine parsing results with the standard parsing results, and obtain the parsing scores for each parsing item. The parsing score is the weight factor.
[0008] Conduct comprehensive analysis using the weight factors. According to the results of the comprehensive analysis, use simulators of multiple different systems and real devices of the corresponding systems for running simulation, label the simulation results, and adjust the weight factors again.
[0009] Furthermore, different parsing items have different algorithms when calculating the weight factors of each engine, specifically including:
[0010] Define file information, application information, and certificate information as the first type of information. The first type of information has fixed results. Directly compare the corresponding engine parsing result x with the standard parsing result y. If x = y, it is determined that the engine parsing is successful; otherwise, it is determined to be a failure. Furthermore, the weight factor of the first type of information = the number of successful parses / the total number of samples.
[0011] Define API call analysis, application permissions, hard-coded key information, source code analysis, and string parsing as the second type of information. The parsing results of the second type of information are greater than or equal to two, and each parsing result is accurate. Furthermore, the weight factor of the second type of information = the total sum of engine parsing result quantities / the total sum of parsing quantities of all samples.
[0012] For the second type of information, there are multiple parsing results. These results do not need to specifically focus on accuracy because basically all the parsed results are accurate, and the more the result content, the better.
[0013] Define the adjustable third-party information, domain name analysis, URL analysis, email analysis, mobile phone number analysis, certificate analysis, manifest analysis, and dynamic packet capture analysis as the third type of information. The parsing results of the third type of information are greater than or equal to two, where there is only one or two pieces of key useful information. Furthermore, assign a trustworthy label to the key useful information in the parsing results, calculate the total sum of the trustworthy label quantities of all samples. For the parsing results, as long as one result is consistent with the value of the trustworthy label, increment the corresponding matching score variable by 1. Furthermore, the weight factor = the matching score variable / the total sum of the trustworthy label quantities of all samples.
[0014] For the third type of information, there are multiple parsing results, but it's not that the more the better, because there are often only one or two key clues involved. Therefore, the more refined and accurate the results are, the better.
[0015] Collect the parsing results of APPs that have been verified through actual combat as samples. These results have been tested through actual combat, and the parsing results are true and reliable.
[0016] From the above several calculations, the weight of each parsing engine for each parsing item can be obtained. When a new application is sent for parsing, the application will be sent to each parsing engine synchronously and in parallel, and different processing will also be carried out for each type of parsing item.
[0017] Furthermore, after calculating the weight factors of each engine, it is also necessary to perform comprehensive analysis using the weight factors, specifically including:
[0018] The first type of information has a fixed result. If only one engine parsing result x is parsed out by multiple engines, directly adopt the engine parsing result x. If the number of engine parsing results parsed out is greater than one, it is necessary to add the weights of multiple engines with the same engine parsing result, and finally adopt the result with the highest weight;
[0019] The number of parsing results of the second type of information is greater than or equal to two. Therefore, take the union of all engine results, and eliminate the parsing engines with a weight < 0.5;
[0020] The number of parsing results of the third type of information is greater than or equal to two. Record each parsing result, add the weights of all engines that have parsed out the corresponding result. If the total weight > 0.6, it is used as the final parsing result.
[0021] Furthermore, the specific operation of running simulations using emulators of multiple different systems and real devices of the corresponding systems includes: using the parsing results of third-party information as input data. If the real devices have obtained parsing conclusions, directly adopt the union of all parsing conclusions of the real devices and discard the parsing conclusions of the emulators. If the real devices have no conclusions, adopt the parsing conclusions of the emulators.
[0022] For static analysis, relying on different engine fusion algorithms can yield relatively accurate conclusions. However, for dynamic analysis, there is also the influence of the operating environment. This method prepares simulators of multiple different systems and real devices of different systems for running simulations to ensure the integrity of the analysis results. At the same time, different processing methods will be adopted for different analysis items.
[0023] Each of the above-obtained analysis results and their union are used as the final result of a single analysis engine. By participating this result in the calculation of the weight factor, a more reliable conclusion can be obtained.
[0024] According to the second aspect of the present invention, a computer program product is proposed, on which one or more computer programs are stored. When the one or more computer programs are executed by a computer processor, the above-mentioned method is implemented.
[0025] One or more of the above technical solutions in the embodiments of the present application have at least one of the following technical effects:
[0026] A SaaS multi-engine parsing method applicable to APP analysis proposed by the present invention can be widely applied in the field of electronic data forensics. It can not only quickly and effectively import APPs for analysis, but also output relatively reliable results through self-comparison and optimization. This method has high reliability and flexibility, and also has strong flexible scalability. The analysis results are structured data and can be imported into different products and systems.
[0027] To help improve Internet governance work, this method has greatly improved the automated analysis results of the dynamic and static behaviors of mobile applications, helping investigators complete the intricate analysis process, saving a lot of time, and improving the efficiency of case detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The drawings illustrate the embodiments and, together with the description, are used to explain the principles of the present invention. Other embodiments and many of the intended advantages of the embodiments will be readily appreciated as they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily to scale with each other. Like reference numerals refer to corresponding similar components.
[0029] Figure 1 A flowchart showing a method for SaaS multi-engine parsing applicable to APP analysis according to an embodiment of the present invention is shown.
[0030] Figure 2 A flowchart showing the process of calculating the weight factor of each engine according to an embodiment of the present invention is shown.
[0031] Figure 3 Shows a schematic diagram of the comprehensive analysis process of weight factors according to an embodiment of the present invention.
[0032] Figure 4 Shows a schematic diagram of the parsing interface of the first type of information according to an embodiment of the present invention.
[0033] Figure 5 Shows a schematic diagram of the parsing interface of the second type of information according to an embodiment of the present invention.
[0034] Figure 6 Shows a schematic diagram of the parsing interface of the third type of information according to an embodiment of the present invention.
[0035] Figure 7 Is a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners
[0036] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the convenience of description, only parts related to the relevant invention are shown in the drawings.
[0037] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0038] Figure 1 Shows a schematic diagram of the process of a method suitable for APP analysis SaaS multi - engine parsing according to an embodiment of the present invention, as Figure 1 shown:
[0039] Use SaaS for distributed system deployment, introduce a parsing engine for parsing to obtain parsing results. Among them, the engine includes: basic application parsing, decompilation analysis, unpacking analysis, virus analysis, website analysis, and script analysis. Calculate the weight factors of each engine using the parsing results, specifically including: collecting the APP parsing results that have been verified in actual combat as samples and importing them into the corresponding engine for parsing, comparing the engine parsing results with the standard parsing results, and obtaining the parsing score of each parsing item. The parsing score is the weight factor;
[0040] In the calculation of the weight factors of each engine, different parsing items have different algorithms. The specific flowchart is as Figure 2 shown, and the steps include:
[0041] Define file information, application information, and certificate information as the first type of information. The first type of information has a fixed result. Directly compare the corresponding engine parsing result x with the standard parsing result y. If x = y, it is determined that the engine parsing is successful; otherwise, it is determined to be a failure. Furthermore, the weight factor of the first type of information = the number of successful parses / the total number of samples;
[0042] The first type of information ensures the accuracy of the engine's parsing of information with fixed results. This type of information usually has a clear standard answer, and the parsing ability of the engine can be quickly evaluated through direct comparison.
[0043] Define API call analysis, application permissions, hard-coded key information, source code analysis, and string parsing as the second type of information. The parsing results of the second type of information are greater than or equal to two, and each parsing result is accurate. Furthermore, the weight factor of the second type of information = the total sum of the engine parsing result quantities / the total sum of the parsing quantities of all samples;
[0044] The second type of information evaluates the engine's parsing ability for multi-result information. This type of information usually involves multiple accurate results. By calculating the total sum of the parsing results, the comprehensiveness and accuracy of the engine in processing multi-result information can be evaluated.
[0045] Define adjustable third-party information, domain name analysis, URL analysis, email analysis, mobile phone number analysis, certificate analysis, manifest analysis, and dynamic packet capture analysis as the third type of information. The parsing results of the third type of information are greater than or equal to two, and there is only one or two pieces of key useful information among them. Furthermore, label the key useful information in the parsing results with a trusted label, calculate the total sum of the trusted label quantities of all samples. For the parsing results, as long as one result is consistent with the value of the trusted label, the corresponding matching score variable is incremented by 1. Furthermore, the weight factor = the matching score variable / the total sum of the trusted label quantities of all samples.
[0046] The third type of information evaluates the engine's recognition ability for key useful information. This type of information usually contains multiple parsing results, but only a few of them are key useful. By labeling the key useful information with a trusted label, the accuracy of the engine in identifying and extracting key information can be evaluated.
[0047] Figure 2Shows the specific steps parsed by the parsing engine M, where three types of information are defined. The first type of information evaluates the parsing accuracy of the fixed result information by the evaluation engine. The second type of information evaluates the comprehensiveness and accuracy of the parsing of multi-result information by the evaluation engine. The third type of information: evaluates the recognition ability of the evaluation engine for key useful information. This step can comprehensively evaluate the performance of the parsing engine on different information types, ensuring its reliability and effectiveness in various application scenarios. These evaluation methods not only help developers understand the advantages and disadvantages of the engine but also can guide the optimization and improvement of the engine.
[0048] Use the weight factor for comprehensive analysis. According to the result of the comprehensive analysis, use the simulators of multiple different systems and the real devices of the corresponding systems for running simulation, label the results of the simulation, and adjust the weight factor again.
[0049] After calculating the weight factors of each engine, it is also necessary to use the weight factor for comprehensive analysis, such as Figure 3 shown in the engine analysis of the new application L, specifically including:
[0050] The first type of information has a fixed result. If only one engine parsing result x is parsed by multiple engines, directly adopt the engine parsing result x. If the number of engine parsing results parsed is greater than one, it is necessary to add the weights of multiple engines with the same engine parsing result, and finally adopt the result with the highest weight;
[0051] In the case of a unique result for the first type of information, directly adopt this result to avoid unnecessary complexity. When there are multiple results, by adding the weights of the same engine parsing result, select the result with the highest weight to ensure the accuracy and reliability of the final result. This method can effectively reduce misjudgment and improve the accuracy of parsing.
[0052] The number of parsing results of the second type of information is greater than or equal to two. Therefore, take the union of all engine results, and among them, eliminate the parsing engines with a weight <0.5;
[0053] Taking the union ensures that all possible parsing results are considered, avoiding omission of important information; by eliminating the engines with a weight lower than 0.5, exclude those less reliable parsing results, and improve the credibility and accuracy of the overall parsing result.
[0054] The number of parsing results of the third type of information is greater than or equal to two. Record each parsing result, add the weights of all engines that parsed the corresponding result. If the total weight >0.6, it is used as the final parsing result.
[0055] The overall credibility of each parsing result is evaluated by adding up the weights of all the engines that have parsed the corresponding results. By setting a total weight threshold (0.6), it is ensured that the parsing results finally adopted have a high credibility. This method can effectively filter out the results with low credibility and improve the accuracy and reliability of the parsing.
[0056] The operation simulation using the simulators of multiple different systems and the real devices of the corresponding systems specifically includes: using the parsing results of third-party information as input data. If the real devices have obtained parsing conclusions, directly take the union of all the parsing conclusions of the real devices and discard the parsing conclusions of the simulators. If the real devices have no conclusions, then adopt the parsing conclusions of the simulators.
[0057] In this embodiment, as Figure 4 shows the interface display of the parsing results of the first type of information, as Figure 5 shows the interface display of the parsing results of the second type of information, as Figure 6 shows the interface display of the parsing results of the third type of information. The three result schematic diagrams show the specific parsing result contents in the business embodiment.
[0058] Since the current algorithm is obtained through training with existing samples, to keep up with the times, more samples and more accurate calculations are needed. Therefore, this method also considers measures to improve accuracy. For the parsed result set, it can directly participate in actual combat tests. If it is found that the results are deviated, then a label can be attached to this parsing conclusion, and all the engines that have parsed this conclusion will adjust their weights again according to the weight factor calculation algorithm. Correspondingly, the engines that have parsed the correct results can also recalculate and increase their weights according to the weight factor calculation algorithm. Thus, the purpose of continuous iteration and continuous improvement is achieved.
[0059] When this method is productized, various mechanisms will also be built, such as clue sharing and point systems, to achieve a virtuous cycle and improve the perfection and accuracy of this method. The more label clues are contributed and the more are enjoyed, the more points will be obtained, attracting more users to join the maintenance, and the higher the accuracy of the engine will be.
[0060] At the same time, new engines will also be introduced regularly, pooling ideas, refining various open-source parsing engines, and adding a variety of self-developed and optimized engines to make the parsing ability more in line with the times.
[0061] Through the SaaS transformation of APP analysis, the reliability of this method is greatly improved, ensuring the timeliness of batch APP analysis. By introducing a multi-engine self-optimization mechanism, it deeply analyzes and intelligently analyzes the APP in the product, and realizes related functions such as static and dynamic analysis, clue extraction, associated line expansion, suspiciousness analysis, and investigation guidance, so as to realize the positioning, analysis, and traceability of the target APP and URL; it can strongly and quickly grasp the key clues hidden in the target APP and URL, and promote the further study of the event.
[0062] Reference is made below to Figure 7 , which shows a schematic structural diagram of a computer system 700 of an electronic device suitable for implementing the embodiments of the present application. Figure 7 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0063] As Figure 7 shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the system 700 are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0064] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that the computer program read from it can be installed into the storage section 708 as needed.
[0065] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above-described functions defined in the methods of the present application are performed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0066] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0067] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0068] The modules described in the embodiments of this application can be implemented in software or in hardware.
[0069] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: deploy a distributed system using SaaS, introduce a parsing engine to obtain a parsing result through parsing, wherein the engine includes: basic application parsing, decompilation analysis, unpacking analysis, virus analysis, website analysis and script analysis, calculate the weight factors of each engine using the parsing result, specifically including: collecting the parsing results of APPs that have been verified in actual combat as samples and importing them into the corresponding engine for parsing, comparing the engine parsing results with the standard parsing results to obtain the parsing scores of each parsing item, and the parsing score is the weight factor; perform comprehensive analysis using the weight factor, and according to the result of the comprehensive analysis, use simulators of multiple different systems and real devices of the corresponding systems to perform operation simulation, label the simulation results and adjust the weight factor again.
[0070] The above description is only a preferred embodiment of the present application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present application.
Claims
1. A method for multi-engine analysis of APP analysis SaaS, characterized in that: include: Utilize SaaS for distributed system deployment, introduce a parsing engine for parsing to obtain parsing results, wherein the engine includes: basic application parsing, decompilation analysis, shelling analysis, virus analysis, URL analysis and script analysis, and use the parsing results to calculate the weight factor of each engine, specifically including: collecting APP parsing results that have been verified in actual combat as samples and importing them into the corresponding engine for parsing, comparing the engine parsing results with the standard parsing results, and obtaining the parsing score of each parsing item, and the parsing score is the weight factor; The weight factors are used to perform a comprehensive analysis, and based on the results of the comprehensive analysis, simulators of multiple different systems and real machines of corresponding systems are used to perform operation simulations, the simulation results are labeled and the weight factors are adjusted again.
2. The method according to claim 1, characterized in that The specific different analysis items in calculating the weight factors of each engine have different algorithms, including: File information, application information and certificate information are defined as the first type of information. The first type of information has a fixed result. The corresponding engine parsing result x is directly compared with the standard parsing result y. If x=y, the engine parsing is determined to be successful, otherwise it is determined to be failed. Then, the weight factor of the first type of information = the number of successful parsing / the total number of samples; API call analysis, application permissions, hard-coded key information, source code analysis, and string parsing are defined as the second type of information. The parsing results of the second type of information are greater than or equal to two, where each parsing result is accurate. The weight factor of the second type of information = the total number of engine parsing results / the total number of parsed results of all samples; The third-party information that can be adjusted, domain name analysis, URL analysis, email analysis, mobile phone number analysis, certificate analysis, manifest analysis and dynamic packet capture analysis are defined as the third category of information. The parsing results of the third category of information are greater than or equal to two, among which there are only one or two key useful information. Then, the key useful information in the parsing results is marked with a trusted label, and the total number of trusted labels of all samples is calculated. For the parsing results, as long as there is one result that is consistent with the value of the trusted label, the corresponding matching score variable is increased by 1, and then the weight factor = matching score variable / the total number of trusted labels of all samples.
3. The method according to any one of claims 1 or 2, characterized in that: After calculating the weight factors of each engine, it is also necessary to use the weight factors to perform a comprehensive analysis, specifically including: The first type of information has a fixed result. If multiple engines parse out only one engine parsing result x, then the engine parsing result x is directly adopted. If there are more than one engine parsing results, then it is necessary to add up the engine weights of the same engine parsing result, and finally adopt the result with the highest weight. The number of parsing results for the second type of information is greater than or equal to two, so the union of all engine results is taken, wherein parsing engines with weights < 0.5 are eliminated; The number of parsing results of the third type of information is greater than or equal to two. Each parsing result is recorded, and the weights of all engines that parse the corresponding results are added together. If the total weight is greater than 0.6, it is used as the final parsing result.
4. The method according to any one of claims 1 or 2, characterized in that: The use of simulators of multiple different systems and real machines of the corresponding systems to perform operation simulation specifically includes: using the analysis results of third-party information as input data, if the real machine has reached an analysis conclusion, directly taking the union of the analysis conclusions of all real machines and discarding the analysis conclusion of the simulator; if the real machine has no conclusion, then using the analysis conclusion of the simulator.
5. A computer program product, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
6. A computing system, characterized in that: The method comprises a processor and a memory, wherein the processor is configured to execute the method according to any one of claims 1 to 4.