Methods, devices, equipment, media, and products for identifying target applications
By extracting stable features of applications and combining them with volatile features for identification and blocking, the problem of traditional identification technologies being unable to keep up with the changing pace of fraudulent apps has been solved, achieving precise and efficient prevention and control, and ensuring user safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE GRP HENAN CO LTD
- Filing Date
- 2026-02-02
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional fraudulent app identification technologies cannot keep up with the rapid changes in the surface features of fraudulent apps, making it difficult to achieve accurate and efficient prevention and control. Furthermore, blocking methods based on static features are ineffective and cannot effectively protect users' property security and personal information rights.
By acquiring network traffic data from applications, extracting their stable features based on the development framework, and using a pre-set stable feature library for identification, the system generates identification results, identifies target applications, and combines this with volatile feature information for blocking, thus overcoming the limitations of static features and achieving continuous prevention and control.
It enables accurate identification and effective blocking even when fraudulent applications frequently change their characteristics, protecting users' property security and personal information rights, avoiding the problem of traditional methods failing due to characteristic changes, and providing continuous and effective prevention and control capabilities.
Smart Images

Figure CN122133131A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, and in particular relates to a method, apparatus, device, medium and product for identifying target applications. Background Technology
[0002] With the rapid development of the mobile internet, fraudulent applications (APPs) have become the main tools for telecommunications and network fraud. They have functions such as remote control, screen sharing, and end-to-end encrypted private messaging. By impersonating legitimate applications, inducing users to disclose sensitive information, and illegally transferring funds, they seriously endanger users' property security and personal information rights.
[0003] Traditional fraudulent app identification technologies rely heavily on manual review, static feature matching, and blacklist blocking. However, because fraudulent apps can evade regulation by frequently changing surface features such as IP addresses, icons, package names, and MD5 values, traditional fraudulent app identification methods cannot keep up with the pace of change and cannot meet the needs of accurate and efficient prevention and control. Furthermore, traditional blocking methods based on static features such as domain names, IP addresses, and MD5 values have become largely ineffective. Summary of the Invention
[0004] This application provides a method, apparatus, device, medium, and product for identifying target applications. Even if fraudulent applications frequently change their surface features, they can be accurately identified based on stable features and effectively blocked through volatile features.
[0005] In a first aspect, embodiments of this application provide a method for identifying a target application, the method comprising: Obtain network traffic data from the application; Based on the network traffic data of the application, extract the stable features of the application, which are the inherent features of the application based on the development framework; The stable features of the application are identified according to a preset stable feature library, and an identification result is generated. The stable feature library includes the stable features of multiple fraudulent application installation package samples and the tags of each fraudulent application installation package sample. The stable features of different fraudulent application installation package samples are different. The identification result is used to indicate whether the stable features of the application match the stable features of any one of the fraudulent application installation package samples. If the stable characteristics of the application identified by the recognition result match the stable characteristics of any of the fraudulent application installation package samples, the application is determined to be the target application. Based on the information of the volatile characteristics of the target application corresponding to the tag in the network traffic data, the target application is blocked. The volatile characteristics are different for different tags.
[0006] Secondly, embodiments of this application provide a target application identification device, the device comprising: The first acquisition module is used to acquire network traffic data of the application. The first extraction module is used to extract stable features of the application based on the network traffic data of the application. The stable features are the inherent features of the application based on the development framework. The identification module is used to identify the stable features of the application according to a preset stable feature library and generate an identification result. The stable feature library includes the stable features of multiple fraudulent application installation package samples and the tags of each fraudulent application installation package sample. The stable features of different fraudulent application installation package samples are different. The identification result is used to indicate whether the stable features of the application match the stable features of any one of the fraudulent application installation package samples. The first determining module is used to determine the application as the target application if the stable characteristics of the application characterized by the identification result match the stable characteristics of any of the fraudulent application installation package samples. The blocking module is used to block the target application based on the information of the volatile characteristics of the tag corresponding to the target application in the network traffic data. The volatile characteristics are different for different tags.
[0007] Thirdly, embodiments of this application provide an electronic device, the device including: a processor and a memory storing computer program instructions; the processor, when executing the computer program instructions, implements the target application identification method as described above.
[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the target application identification method as described in any of the above claims.
[0009] Fifthly, embodiments of this application provide a computer program product, wherein instructions in the computer program product, when executed by a processor of an electronic device, cause the electronic device to perform the target application identification method as described in any of the above.
[0010] The target application identification method, apparatus, device, medium, and product of this application embodiment are capable of acquiring network traffic data of the application; extracting stable features of the application based on the network traffic data, wherein stable features are inherent features of the application based on the development framework; identifying the stable features of the application according to a preset stable feature library, generating an identification result, wherein the stable feature library includes stable features of multiple fraudulent application installation package samples and tags of each fraudulent application installation package sample, wherein the stable features of different fraudulent application installation package samples are different, and the identification result is used to indicate whether the stable features of the application match the stable features of any fraudulent application installation package sample; if the identification result indicates that the stable features of the application match the stable features of any fraudulent application installation package sample, the application is determined to be a target application; and blocking the target application based on the information of the volatile features of the tags corresponding to the target application in the network traffic data, wherein the volatile features corresponding to different tags are different. Thus, in this embodiment, since stable features are inherent to the application based on the development framework and are not easily changed, fraudulent applications can be accurately and efficiently identified in network traffic data based on stable features in the stable feature library, overcoming the limitations of static feature identification. Then, targeted blocking is implemented based on the volatile feature information of its corresponding tag. Even if the fraudulent application updates its volatile features, the blocking strategy can still be re-identified and updated through stable features, avoiding the failure of traditional blocking methods based on static features due to feature changes. This achieves continuous and effective prevention and control of fraudulent applications and effectively protects users' property security and personal information rights. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating a method for identifying a target application provided in an embodiment of this application; Figure 2 This is a flowchart illustrating another method for identifying a target application provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of the target application identification device provided in the embodiments of this application; Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0013] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0014] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0015] With the rapid development of the mobile internet, fraudulent applications (APPs) have become the main tools for telecommunications and network fraud. They have functions such as remote control, screen sharing, and end-to-end encrypted private messaging. By impersonating legitimate applications, inducing users to disclose sensitive information, and illegally transferring funds, they seriously endanger users' property security and personal information rights.
[0016] Traditional fraudulent app identification technologies rely heavily on manual review, static feature matching, and blacklist blocking. However, because fraudulent apps can evade regulation by frequently changing surface features such as IP addresses, icons, package names, and MD5 values, traditional fraudulent app identification methods cannot keep up with the pace of change and cannot meet the needs of accurate and efficient prevention and control. Furthermore, traditional blocking methods based on static features such as domain names, IP addresses, and MD5 values have become largely ineffective.
[0017] The acquisition, storage, use, and processing of data in this application comply with relevant national laws and regulations. It should be noted that certain software, components, models, and other existing industry solutions may be mentioned in the embodiments of this application. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.
[0018] To address the problems of the prior art, embodiments of this application provide a method, apparatus, device, medium, and product for identifying target applications. The method for identifying target applications provided in this application embodiment will be described first below.
[0019] Figure 1 A flowchart illustrating a target application identification method according to an embodiment of this application is shown. Figure 1 As shown, a method for identifying a target application may include the following steps S101 to S105: S101. Obtain network traffic data for the application; S102. Based on the network traffic data of the application, extract the stable features of the application. The stable features are the inherent features of the application based on the development framework. S103. Identify the stable features of the application according to the preset stable feature library and generate identification results. The stable feature library includes the stable features of multiple fraudulent application installation package samples and the labels of each fraudulent application installation package sample. The stable features of different fraudulent application installation package samples are different. The identification results are used to indicate whether the stable features of the application match the stable features of any fraudulent application installation package sample. S104. If the stable characteristics of the application represented by the identification result match the stable characteristics of any fraudulent application installation package sample, the application is identified as the target application. S105. Based on the information of the volatile characteristics of the target application corresponding to the tag in the network traffic data, the target application is blocked. Different tags correspond to different volatile characteristics.
[0020] The target application identification method of this application embodiment can acquire network traffic data of the application; extract stable features of the application based on the network traffic data, the stable features being inherent features of the application based on its development framework; identify the stable features of the application according to a preset stable feature library, generating an identification result, the stable feature library including the stable features of multiple fraudulent application installation package samples and the tags of each fraudulent application installation package sample, the stable features of different fraudulent application installation package samples being different, the identification result being used to indicate whether the stable features of the application match the stable features of any fraudulent application installation package sample; if the identification result indicates that the stable features of the application match the stable features of any fraudulent application installation package sample, the application is determined to be the target application; and the target application is blocked according to the information of the volatile features corresponding to the tags of the target application in the network traffic data, the volatile features corresponding to different tags being different. Thus, in this embodiment, since stable features are inherent to the application based on the development framework and are not easily changed, fraudulent applications can be accurately and efficiently identified in network traffic data based on stable features in the stable feature library, overcoming the limitations of static feature identification. Then, targeted blocking is implemented based on the volatile feature information of its corresponding tag. Even if the fraudulent application updates its volatile features, the blocking strategy can still be re-identified and updated through stable features, avoiding the failure of traditional blocking methods based on static features due to feature changes. This achieves continuous and effective prevention and control of fraudulent applications and effectively protects users' property security and personal information rights.
[0021] In S101, the aforementioned network traffic data may include, but is not limited to, metadata, message length, time series, byte distribution, and unencrypted TLS header information based on Transport Layer Security (TLS) traffic.
[0022] In some embodiments of this application, network traffic data of the application is acquired, for example, by the operator or the terminal side in real time. Specifically, the operator can deploy optical acquisition equipment at the network edge node, acquire network traffic data through deep packet inspection technology, and then use a time window sampling strategy to ensure processing efficiency, while recording five-tuple information to bind user sessions.
[0023] In S102, the aforementioned stable feature can be an inherent feature of the application based on the development framework. For example, it can be a fingerprint feature composed of the specific TCP / UDP traffic generated by the APP during use, its packet length, and the string characteristics of the traffic. This is a deep feature of the APP during use. This fingerprint feature is an inherent feature based on the underlying feature of the development framework, so this feature is very stable and will not change for a considerable period of time.
[0024] In some embodiments of this application, based on the network traffic data of the application, stable features of the application are extracted. For example, this may involve parsing the plaintext header information of the Transport Layer Security (TLS) handshake phase in the network traffic data of the application to extract TLS metadata; extracting message sequence features from the network traffic data of the application; extracting time dimension features from the network traffic data of the application; extracting load features from the network traffic data of the application; and aggregating the TLS metadata, message sequence features, time dimension features, and load features to obtain the stable features of the application.
[0025] In S103, the aforementioned stable feature library may include stable features of multiple fraudulent application installation package samples and tags for each fraudulent application installation package sample. Different fraudulent application installation package samples have different stable features. It should be noted that fraudulent application installation package samples with different tags correspond to different volatile features.
[0026] The above identification results can be used to indicate whether the stability characteristics of an application match the stability characteristics of any sample of a fraudulent application installation package.
[0027] In some embodiments of this application, the stability features of an application are identified according to a preset stability feature library, and an identification result is generated. For example, the stability features of the application can be searched in the stability feature library according to the application's application identifier, software development kit version, and network environment type to locate a target fraudulent application installation package sample that matches the application. The stability features of the application are then identified according to the target fraudulent application installation package sample, and an identification result is generated. The identification result is used to indicate whether the stability features of the application match the stability features of the target fraudulent application installation package sample.
[0028] In S104, since the stable feature library is constructed from the stable features of multiple fraudulent application installation package samples, when an application matches the stable feature of any fraudulent application installation package sample in the stable feature library, the application is identified as the target application, i.e., a fraudulent APP.
[0029] In S105, the aforementioned volatile characteristics, exemplarily, may include, but are not limited to, surface characteristics such as the domain name used, the IP address of the communication server, the icon, the packet name, and the MD5 value. The focus is on the domain name, the IP address of the communication server, and the server IP address port. It should be noted that different tags correspond to different volatile characteristics.
[0030] In some embodiments of this application, the target application is blocked based on the volatile characteristics of the tag corresponding to the target application in network traffic data. For example, this can be achieved by determining the tags of fraudulent application installation package samples matching the target application in a stable feature library as the target tags of the target application; determining the volatile characteristics corresponding to the target tags as the target volatile characteristics of the target application; extracting the information of the target volatile characteristics from the network traffic data; and blocking the target application based on the information of the target volatile characteristics. Alternatively, the tags of fraudulent application installation package samples matching the target application in a stable feature library can be determined as the target tags of the target application; determining the volatile characteristics corresponding to the target tags as the target volatile characteristics of the target application; extracting the information of the target volatile characteristics from the network traffic data; cleaning the information of the target volatile characteristics to obtain standard information of the target volatile characteristics; and blocking the target application based on the standard information of the target volatile characteristics.
[0031] In some embodiments, the above-described S102 may specifically include: Parse the plaintext header information of the transport layer security protocol handshake phase in the network traffic data of the application to extract the transport layer security protocol metadata; Extract packet sequence characteristics from the application's network traffic data; Extracting the time dimension features from the application's network traffic data; Extract load characteristics from the application's network traffic data; By aggregating transport layer security protocol metadata, message sequence characteristics, time dimension characteristics, and payload characteristics, stable characteristics of the application are obtained.
[0032] The aforementioned transport layer security protocol metadata may include the transport layer security protocol version, a list of cipher suites, extension types, and certificate information. The certificate information includes certificate chain length, organizational structure characteristics, certificate validity period, public key algorithm type, and the distribution of Subject Alternate Names (SANs).
[0033] In some embodiments of this application, plaintext header information of the Transport Layer Security (TLS) handshake phase in the network traffic data of the application is parsed to extract TLS metadata. For example, this may involve parsing the plaintext header information of the TLS handshake phase to extract the TLS version, cipher suite list, extension type (SNI / ALPN, etc.) from ClientHello; capturing the cipher suite and certificate chain characteristics selected by ServerHello; recording the Session Ticket length and OCSP binding status; and analyzing the non-encrypted characteristics of the certificate exchange phase: certificate chain length and organizational structure characteristics; certificate validity period and public key algorithm type; and the distribution of the number of Alternate Names (SANs).
[0034] The aforementioned message sequence features may include packet length sequence, packet length distribution, and packet length direction transition matrix.
[0035] In some embodiments of this application, extracting packet sequence features from the network traffic data of an application can, for example, involve extracting the length sequence of the first N data packets within a single stream of the application's network traffic data (distinguishing between uplink and downlink directions), calculating the statistical characteristics of the packet length distribution (mean / variance / skewness / kurtosis), and generating a packet length direction transition matrix (uplink → downlink probability).
[0036] The aforementioned time-dimensional characteristics may include packet arrival time intervals, traffic burst patterns, and the ratio of active to quiet periods within a flow.
[0037] In some embodiments of this application, extracting packet sequence features from the network traffic data of an application can, for example, involve measuring the statistical distribution of packet arrival time intervals (IPD) based on the network traffic data of the application, identifying traffic burst patterns (burst duration / number of packets), and calculating the ratio of active to inactive periods within the flow's lifecycle.
[0038] The aforementioned payload characteristics may include the histogram of the first message payload bytes, the entropy value of the encrypted text, and the padding pattern of the transport layer security protocol record.
[0039] In some embodiments of this application, load characteristics are extracted from the network traffic data of the application. For example, this can be done by constructing a histogram of the first message load bytes based on the network traffic data of the application, analyzing the entropy characteristics of the encrypted text, and detecting the TLS record padding pattern.
[0040] In some embodiments of this application, transport layer security protocol metadata, message sequence features, time dimension features, and payload features are aggregated to obtain stable features of the application. For example, multi-stream features can be aggregated to form APP-level fingerprint features. In this case, TLS metadata is statistically analyzed using the mode (e.g., the most common cipher suites), the time dimension features are calculated as the cross-stream average, and the packet length distribution is weighted and fused to generate a lightweight fingerprint digest. For example, the fingerprint features of a certain payment APP include: TLS 1.3 accounting for 98%, ALPN=h2, packet length sequence pattern [1420→52→1420], and IPD average of 16ms.
[0041] In this embodiment, multi-dimensional feature aggregation improves the comprehensiveness and discriminativeness of stable features, enhancing recognition accuracy; features are strongly bound to the underlying layer, possessing strong anti-mutation capabilities and breaking through the limitations of traditional recognition; multiple features complement each other to cancel interference, adapting to multiple scenarios and improving feature reliability and robustness; thereby providing a highly reliable basis for the closed loop of target application identification, extraction, and blocking, ensuring the effectiveness and timeliness of prevention and control.
[0042] In some embodiments, the aforementioned message sequence features may include packet length sequence, packet length distribution, and packet length direction transition matrix. Before aggregating the transport layer security protocol metadata, message sequence features, time dimension features, and payload features to obtain the stable features of the application, the above method may further include: The packet-length sequence is enhanced to obtain an enhanced packet-length sequence. The enhancement process includes at least one of discrete Fourier transform, temporal symbolic compression, and state transition probability quantization. Normalize the time dimension features to obtain enhanced time dimension features; The above aggregation of transport layer security protocol metadata, message sequence characteristics, time dimension characteristics, and payload characteristics yields stable characteristics of the application, which may specifically include: By aggregating transport layer security protocol metadata, enhanced packet length sequences, packet length distributions, packet length direction transition matrices, enhanced time dimension features, and load features, stable characteristics of the application are obtained.
[0043] The aforementioned enhancement processes include at least one of discrete Fourier transform, temporal symbolic compression, and state transition probability quantization.
[0044] In some embodiments of this application, the packet-length sequence is enhanced to obtain an enhanced packet-length sequence. For example, the packet-length sequence may be subjected to discrete Fourier transform to extract frequency domain features, the packet-length sequence may be compressed using the temporal symbolization (SAX) method, and a Markov model may be constructed to describe the state transition probability.
[0045] In this embodiment, time-domain noise is filtered and core periodic patterns are extracted using discrete Fourier transform. Time-series symbolic compression achieves dimensionality reduction and standardization, and state transition probability quantization transforms static sequences into dynamic regularity features. This effectively solves the problems of high noise, high dimensionality, and shallow pattern characterization in the original packet length sequence, making packet length-related features more closely aligned with the underlying interaction essence of the application, significantly enhancing distinguishability and anti-spoofing capabilities. Furthermore, the impact of network jitter on time-dimensional features is normalized to eliminate feature fluctuations caused by different network environments (WiFi / 4G / 5G) and device states, ensuring consistency of time-dimensional features across scenarios and avoiding recognition bias caused by environmental interference, thus improving the robustness of stable features. In this way, the enhanced packet length sequence and time-dimensional features complement TLS metadata, packet length distribution, packet length direction transition matrix, and load features, retaining core information in each dimension while avoiding the inherent defects of the original features. The aggregated stable features are more compact, accurate, and resistant to interference, adaptable to large-scale real-time traffic identification and fraudulent APP variant prevention scenarios.
[0046] As one implementation of this application, in order to construct a stable feature library, the method may further include the following before S103: Obtained multiple sample installation packages of fraudulent applications; Extract stable features from the installation package samples of each fraudulent application; Based on the stable characteristics of each fraudulent application installation package sample, the label of each fraudulent application installation package sample is determined. Different labels of fraudulent application installation package samples correspond to different volatile characteristics. A stable feature library was constructed based on multiple fraudulent application installation package samples, their stable features, and tags.
[0047] In some embodiments of this application, the acquisition of multiple fraudulent application installation package samples can, exemplarily, involve installing and simulating the use of the app on a test terminal, and then obtaining its traffic data using packet capture software. On the terminal side: the Android platform utilizes VPNService to implement transparent traffic proxying, isolating traffic from different apps through a UID binding mechanism; the iOS platform develops a customized network plugin based on the Network Extension framework to capture raw packets at the data link layer. The terminal acquisition module automatically marks the network environment (WiFi / 4G / 5G) and device status (foreground / background).
[0048] In some embodiments of this application, the stability features of each fraudulent application installation package sample are extracted. The specific implementation of extracting the stability features of the application in S102 above can be referred to, and will not be repeated here.
[0049] The tags for the aforementioned fraudulent application installation package samples can, for example, be categorized and bound by fraud family (such as XX remote control family, XX order-brushing family), fraud type (such as fake investment, screen sharing, telecommunications fraud), and core malicious SDK version (such as MalSDK_v2.1, FraudSDK_v3.0).
[0050] The aforementioned fraudulent application installation package samples with different tags correspond to different volatile characteristics. For example, samples with the same tag share the same set of volatile characteristics. Fraudulent apps from the same fraudulent family, the same malicious SDK version, or the same type of fraud have highly similar communication strategies and evasion methods due to their consistent underlying development logic, and thus correspond to the same set of volatile characteristics (e.g., the XX remote control family tag corresponds to its exclusive remote control server IP segment, communication port 21116, and exclusive remote control domain name prefix). The volatile characteristic sets of samples with different tags are significantly different. Fraudulent apps from different fraudulent families and with different malicious SDK versions have completely different communication servers, communication ports, and domain name rules, and thus correspond to different sets of volatile characteristics (e.g., the tag for fraudulent apps involving fake orders corresponds to e-commerce counterfeit domain names and port 8080, while the tag for fraudulent apps involving remote control corresponds to overseas server IPs and port 21116).
[0051] In some embodiments of this application, a stable feature library is constructed based on multiple fraudulent application installation package samples, their stable features, and tags. For example, this can be achieved by storing the multiple fraudulent application installation package samples, their stable features, and tags according to a three-tiered index architecture based on application identifier, software development kit version, and network environment type. Alternatively, a two-tiered associative storage structure can be used, forming a three-level linkage with the stable feature library and the fraudulent APP sample library: First, the stable feature library connects to the fraudulent APP installation package samples and then to the sample tags, achieving a precise mapping between samples and tags; second, the tag and volatile feature association library connects to the tag index and then to a dedicated volatile feature set (including feature values, feature types, and update times), achieving a unique mapping between tags and volatile features.
[0052] In this embodiment, by batch acquiring samples of fraudulent APP installation packages, extracting stable features, assigning differentiated labels to different volatile features according to the features, and constructing a stable feature library, a highly reliable data foundation for identifying fraudulent APP variants is laid, enabling accurate classification and differentiated blocking management of fraudulent APPs. Furthermore, an efficient feature retrieval and matching system is constructed to adapt to large-scale real-time identification needs. Simultaneously, the scalability and timeliness of the identification and blocking closed loop are enhanced, effectively addressing scenarios of rapid mutation of fraudulent APPs. It also improves the ability to identify fraudulent APPs within the same family, expanding the scope of prevention and control, and providing core underlying support for accurate, efficient, and comprehensive prevention and control of fraudulent APPs.
[0053] In some embodiments, the stable feature library constructed based on multiple fraudulent application installation package samples, the stable features and tags of each fraudulent application installation package sample, may specifically include: A stable feature library is constructed by storing multiple fraudulent application installation package samples, stable features and tags of each fraudulent application installation package sample according to a three-level index architecture of application identifier, software development kit version and network environment type. The aforementioned network traffic data may include application identifiers, software development kit versions, and network environment types. Specifically, S103 may include: Based on the application's application identifier, software development kit version, and network environment type, a search is conducted in the stable feature database to locate target fraudulent application installation package samples that match the application. Based on the target fraudulent application installation package sample, the stability characteristics of the application are identified, and the identification results are generated. The identification results are used to indicate whether the stability characteristics of the application match the stability characteristics of the target fraudulent application installation package sample.
[0054] In some embodiments of this application, multiple fraudulent application installation package samples, their stable features and tags are stored according to a three-layer index architecture of application identifier, software development kit version, and network environment type to construct a stable feature library. For example, multiple fraudulent application installation package samples, their stable features and tags are stored using a three-layer index structure: APP identifier to SDK version to network environment type. Each APP stores multiple version feature snapshots to support version evolution tracking. The feature vectors are stored using an efficient binary format (such as HDF5) to construct a stable feature library.
[0055] The aforementioned sample installation packages of fraudulent applications can be found in a stable feature library that match the application's installation package.
[0056] In some embodiments of this application, a search is performed in a stable feature database based on the application's application identifier, software development kit version, and network environment type to locate a target fraudulent application installation package sample that matches the application. For example, the stable feature database can be searched according to a three-layer index architecture of application identifier, software development kit version, and network environment type to locate a target fraudulent application installation package sample that matches the application.
[0057] The above identification results can be used to indicate whether the stability characteristics of the application match the stability characteristics of the target fraudulent application installation package sample.
[0058] In this embodiment, a stable feature library is constructed by using a three-layer index architecture of APP identifier, SDK version, and network environment type to collect samples of fraudulent APP installation packages, stable features, and tags. During identification, the target samples are accurately retrieved and located based on this three-layer index, and stable feature matching is completed. This effectively improves feature retrieval efficiency and identification accuracy, enhances the structured management and version evolution tracking capabilities of the feature library, and adapts to the identification needs of fraudulent APPs in multiple network environments, providing efficient and reliable feature matching support for subsequent accurate blocking.
[0059] In some embodiments, the above-described S105 may specifically include: The tags of fraudulent application installation package samples that match the target application in the stable feature library are identified as the target tags of the target application. The volatile features corresponding to the target label are identified as the target volatile features of the target application. Extracting information about the target's volatile characteristics from network traffic data; Based on information about the target's volatile characteristics, the target application is blocked.
[0060] The aforementioned target tags can be the tags of fraudulent application installation package samples that match the target application in the stable feature library.
[0061] The aforementioned target volatility features can be the volatility features corresponding to the target label.
[0062] For example, the information regarding the aforementioned volatile characteristics of the target may include the domain name, the IP address of the communication server, and specific information about the server's IP address and port.
[0063] In this embodiment, the target application is blocked based on information about its volatile characteristics. For example, this can involve real-time identification and immediate blocking of fraudulent apps and their enabled communication server IP addresses. The communication server IP address can be blocked within the same operator, or it can be shared with internet companies, terminal manufacturers, or other operators for joint blocking. Blocking methods can utilize the operator's DNS, DPI, and CMNET blackhole routing, etc.
[0064] In this embodiment, the target tag of the target application is determined by matching a stable feature library, and then the corresponding target volatile features are obtained and its network traffic information is extracted for blocking. This achieves precise linkage from stable feature identification to volatile feature blocking, effectively improving the targeting and timeliness of blocking, avoiding the problem of traditional static feature blocking being prone to failure. At the same time, relying on the association between tags and volatile features, it can quickly adapt to the updates of volatile features of fraudulent APPs, forming an efficient closed loop of identification, extraction and blocking, and enhancing the continuous prevention and control capabilities against variants of fraudulent APPs.
[0065] In some embodiments, blocking the target application based on the information of the target's volatile characteristics may specifically include: The information on the target's volatile features is cleaned and processed to obtain standard information on the target's volatile features; Based on standard information about the target's volatile characteristics, the target application is blocked.
[0066] In some embodiments of this application, the above-mentioned cleaning process may include at least one of deduplication and noise reduction, validity verification, format standardization, and redundancy removal.
[0067] The deduplication process involves removing duplicate and redundant information. A deduplication benchmark is established based on the core attributes of the target's volatile characteristics to eliminate the interference of duplicate data on the accuracy of blocking. Specific rules are as follows: Network connectivity features: Using IP address and domain name as the primary key, and port and DNS resolution records as secondary keys, completely duplicate entries are removed. For scenarios where the same domain name corresponds to multiple mirror IPs, IPs with normal activity are retained and merged into an IP segment set, while invalid mirror addresses are removed. Transmission link and application surface layer features: Using proxy node address and temporary download link as unique identifiers, duplicate information with completely identical content or only character redundancy (such as spaces and special characters) is removed, retaining one standard entry. After deduplication, a preliminary set of volatile features without duplicates is generated, and the frequency of each feature is marked to provide a reference for subsequent effectiveness judgment.
[0068] Validity verification filters out invalid and abnormal information. Through multi-dimensional verification rules, it eliminates invalid, forged, and mismatched features with the target application, ensuring that retained information has actual blocking value: Liveness verification: For network accessibility features such as IP addresses, domain names, and proxy nodes, it uses ping testing, port connectivity testing, and domain name resolution validity verification to mark and eliminate features that are unreachable or have failed resolution. Attribute matching verification: Combining the attributes of the fraudulent app corresponding to the target tag (SDK version, fraud family), it verifies whether volatile features conform to the communication patterns of this type of app (such as specific fraud family-specific ports, domain name prefixes), filtering out abnormal features with mismatched attributes (such as ports not commonly used by the target family). Legality verification: Through regular expression matching, it eliminates forged features (such as incorrectly formatted IPs, invalid domain names without suffixes, and maliciously concatenated fake links), retaining features that conform to network protocol specifications and have actual communication / access capabilities.
[0069] Standardized formatting and unified feature representation specifications ensure that valid feature information is uniformly converted according to preset format specifications, guaranteeing standardized representation of volatile features from different sources and in different formats, adapting to the parsing and calling requirements of subsequent blocking devices. For network addresses: IP addresses are uniformly converted to IPv4 / IPv6 standard format, removing prefix spaces and redundant comment characters; domain names are uniformly converted to lowercase format, removing prefixes such as "http: / / " and "https: / / ", retaining the core domain name and suffix, e.g., standardizing "HTTPS: / / ABC.COM / " to "abc.com". For ports and links: port numbers are uniformly retained in numeric format, removing redundant characters in port range descriptions; proxy node and tunnel addresses are integrated according to a fixed format of "protocol type + address + port", e.g., "socks5: / / 192.168.1.1:1080". For application surface layers: temporary download links have redundant parameters removed, retaining the core link address; dynamic icon identifiers are standardized according to platform type and identifier value format to ensure consistency across terminals.
[0070] Noise reduction and redundancy removal streamline feature information by eliminating redundant fields and interfering information irrelevant to blocking, thus simplifying the feature set and improving the efficiency of subsequent blocking operations. Redundant fields are removed by deleting log notes, traffic collection timestamps, and non-core parameters such as statistical parameters in links and geographical location notes corresponding to IP addresses, retaining only the core information necessary for blocking. Interference filtering filters out fragmented features caused by network fluctuations and collection errors, such as invalid links that are too short and randomly generated temporary character identifiers, retaining core features that are strongly related to communication and operation of the target application.
[0071] In this embodiment, the information of the target's volatile features is first cleaned to obtain highly accurate standard information of the target's volatile features, and then the target application is blocked based on this information, which can further improve the accuracy of blocking the target application.
[0072] As another implementation of this application, in order to reduce the risk of false blocking, before S105 above, the method may further include: Obtain the risk level of at least one user using the target application; Determine the risk level of the target application based on the risk level of at least one user; Specifically, S105 mentioned above may include: If the risk level of the target application is higher than the preset risk target level, the target application will be blocked based on the information of the target's volatile characteristics.
[0073] In some embodiments of this application, the risk level of at least one user using the target application is obtained, for example, by (1) user group positioning: by associating and locking at least one user using the target application through the application identifier and user device identifier (such as IMEI, UUID) in the network traffic data, forming a user list; for the target application with a large number of users, a stratified sampling strategy is adopted to select sample users (the sample size is not less than 50, covering different network environments and device types) to ensure the representativeness of the assessment. (2) individual user risk level determination: based on the risk factors and weights prepared in advance, each user is quantitatively scored (total score 10 points), and the risk level is determined according to the score: 8-10 points corresponds to level 1 risk, 6-7.9 points corresponds to level 2 risk, 4-5.9 points corresponds to level 3 risk, and 0-3.9 points corresponds to level 4 risk; at the same time, the anti-fraud system history is retrieved, and users with clear fraud-related marks are directly determined as level 1 risk without repeated scoring. (3) data verification and deduplication: invalid user data (such as users with duplicate device identifiers or more than 30% missing data) are removed to ensure that the risk level determination of each user is based on sufficient evidence and the data is accurate.
[0074] In some embodiments of this application, the risk level of a target application is determined based on the risk level of at least one user. For example, a combined algorithm of weighted average and level percentage correction can be used. This algorithm combines user risk levels and weights to calculate the comprehensive risk score of the target application, thus determining the application's risk level: Step 1: Weighted score calculation. The formula is: Comprehensive score = Σ (Individual user risk score × Corresponding user weight) / Total number of users; where the individual user risk score is assigned according to level (Level 1: 10 points, Level 2: 7 points, Level 3: 5 points, Level 4: 2 points). Step 2: Level percentage correction. If the proportion of Level 1 / Level 2 risk users is ≥30%, the comprehensive score increases by 10%; if the proportion of Level 4 risk users is ≥70%, the comprehensive score decreases by 10%. The corrected score corresponds to the application risk level threshold (8-10 points for Level 1, 6-7.9 points for Level 2, 4-5.9 points for Level 3, 0-3.9 points for Level 4). Special scenario determination: If there is at least one Level 1 risk user, and that user suffers clear fraud losses after using the target application, the target application is directly determined to be Level 1 risk, without needing to calculate the comprehensive score.
[0075] In this embodiment, user risk level aggregation is introduced to determine the application risk level before blocking. Only target applications with risk levels higher than a preset threshold are blocked based on the volatile feature information of the corresponding tags. This achieves differentiated and precise prevention and control, improves the targeting and resource utilization of blocking, reduces the risk of false blocking, improves the closed loop of identification, assessment and blocking, enhances the dynamic response capability and the flexibility of the prevention and control system against fraudulent APP variants, and provides reliable support for efficient, controllable and precise prevention and control of fraudulent APPs.
[0076] To facilitate understanding of the target application identification method in the embodiments of this application, the actual application process of this target application identification method is described as follows: like Figure 2 As shown, the overall process of this application embodiment mainly includes the following steps: Step S201: Obtain a sample installation package of the fraudulent APP. Sample files of fraudulent installation packages were extracted from the phones of victims or suspects. These packages, which included but were not limited to those with features such as remote control, screen sharing, and end-to-end encrypted private messaging, were obtained daily from the police.
[0077] Step S202: Extract stable features of the APP Analyze the development framework and / or SDK features of the installation package. Obtain the app's behavioral characteristics through installation and simulated use, such as the specific TCP / UDP traffic generated during app usage and communication with specific server IP addresses.
[0078] The characteristics of Android application package (APK) files are defined into two main categories: volatile characteristics and stable characteristics.
[0079] The volatile characteristics of an app include, but are not limited to, its domain name, server IP address, icon, package name, MD5 value, and other surface-level features. The focus is on the domain name, server IP address, and server IP address port.
[0080] The IP address of the communication server for fraudulent apps is the IP address of the servers hosting the app. After the app connects to the internet, it communicates with this server's IP address. Currently, most server resources are deployed in public clouds, offering great flexibility. Therefore, if an IP address is blocked, fraudsters will immediately use a new address, and they will also maintain a large reserve of server address resources. Fraudsters only need to replace the blocked address with the new communication server IP address in the installation package file and republish it on the internet (the download link for this new address is constantly updated). This allows them to bypass the blocking and continue using the app. Although the communication server IP address of the UnionPay Meeting app is not critical, it changes very rapidly.
[0081] The IP address of the communication server can also be obtained through domain name resolution.
[0082] An app's stable characteristics include, but are not limited to, its fingerprint features. During use, an app generates specific TCP / UDP traffic; the packet length and string characteristics of this traffic constitute a fingerprint for identification. This is a deep-level characteristic of the app during use. This fingerprint feature is based on the underlying characteristics of the development framework, so it is very stable and will not change over a considerable period. The specific implementation scheme for obtaining the app's stable characteristics is as follows.
[0083] 1. Obtain traffic data for a specific app The app was installed and simulated on a test terminal, and its traffic data was obtained using packet capture software. On the terminal side: the Android platform utilizes VPNService to implement transparent traffic proxying, isolating traffic from different apps through a UID binding mechanism; the iOS platform uses a customized network plugin developed based on the Network Extension framework to capture raw packets at the data link layer. The terminal acquisition module automatically marks the network environment (WiFi / 4G / 5G) and device status (foreground / background).
[0084] 2. TLS metadata extraction Parsing the plaintext header information during the TLS handshake phase: Extract the TLS version, cipher suite list, and extension type (SNI / ALPN, etc.) from ClientHello. Capture the characteristics of the selected cipher suite and certificate chain of ServerHello; Record the Session Ticket length and OCSP binding status; Analyze the unencrypted features of the certificate exchange phase: Certificate chain length and organizational structure characteristics; Certificate validity period and public key algorithm type; Distribution of the number of Alternate Names (SANs).
[0085] 3. Dynamic Flow Feature Extraction Message sequence characteristics: Extract the length sequence of the first N data packets in a single stream (distinguishing between uplink and downlink directions), calculate the statistical characteristics of the packet length distribution (mean / variance / skewness / kurtosis), and generate a packet length direction transition matrix (uplink → downlink probability).
[0086] Time dimension features: Statistical distribution of packet arrival time interval (IPD) is measured to identify traffic burst patterns (burst duration / number of packets) and calculate the ratio of active / quiet periods within the flow's lifecycle.
[0087] Payload characteristics: Construct a histogram of the first message payload bytes, analyze the entropy characteristics of the encrypted text, and detect TLS record padding patterns.
[0088] 4. Feature enhancement techniques Frequency domain features are extracted from packet-long sequences using Discrete Fourier Transform (DFT). The SAX (Synchronous Symbolization of Time) method is used to compress the packet-long sequences, and a Markov model is constructed to describe the state transition probabilities. Temporal feature normalization is performed to address network jitter.
[0089] 5. Feature Storage Architecture A three-layer index structure is adopted: APP identifier → SDK version → network environment type. Each APP stores multiple version feature snapshots, supports version evolution tracking, and uses an efficient binary format to store feature vectors (such as HDF5).
[0090] 6. Fingerprint Feature Construction Multi-stream features are aggregated to form an application-level fingerprint. TLS metadata is analyzed using mode statistics (e.g., the most common cipher suites), time-dimensional features are averaged across streams, and packet length distributions are weighted and fused to generate a lightweight fingerprint digest. Example: A payment app's fingerprint contains: TLS 1.3 accounting for 98%, ALPN=h2, packet length sequence pattern [1420→52→1420], and an average IPD of 16ms.
[0091] Step S203: Identify the APP based on stable characteristics in network traffic. Operators or terminal devices acquire network traffic data in real time. Operators can deploy optical sampling devices at network edge nodes to acquire traffic through deep packet inspection technology. A time window sampling strategy is used to ensure processing efficiency, while simultaneously recording five-tuple information to bind to user sessions.
[0092] Based on the acquired traffic, including but not limited to metadata, message length, time series, byte distribution, and unencrypted TLS header information, features are extracted and calculated. These features are then matched with stable features such as the fingerprint feature library of a specific APP for identification, and corresponding tags are assigned.
[0093] Step S204: Extract the volatile features corresponding to the APP traffic data. Based on the identified app tags, extract the volatile characteristics of the app's traffic from the network logs, such as the IP address of the communication server, port, and IP address attribution. This information can be extracted in real time from massive amounts of logs. Network log traffic data generally includes information such as source address, source port, destination IP address, and destination port. Among these, the destination IP address and destination port correspond to the volatile characteristics of the app.
[0094] Further data cleaning can be performed to ensure accuracy. Based on the app's communication server IP address, its location and port, and its historical usage, fraudulent apps can be further identified, and the corresponding communication server IP addresses can be extracted. Preferably, apps meeting the following characteristics can be further confirmed as fraudulent: some apps have relatively fixed port numbers, their IP addresses are located in certain overseas regions, and they are first-time users of a particular victim (screen-sharing fraudulent apps).
[0095] Step S205: Blocking on the network based on volatile characteristics. It can identify fraudulent apps and their enabled communication server IP addresses in real time and block them immediately. Blocking of communication server IP addresses can be done within the same operator or shared with internet companies, terminal manufacturers, and other operators for joint blocking. Blocking methods can utilize the operator's DNS, DPI, and CMNET blackhole routing, among others.
[0096] Furthermore, based on the user's risk level, the risk of fraudulent apps and IP addresses is confirmed. For some victims using fraudulent apps, the user's risk level is comprehensively assessed by combining various risky behaviors, and data is output in real time for dissuasion. At this time, based on the user's risk level and the dissuasion efforts, the risk and accuracy of the identification of fraudulent apps and IP addresses are confirmed. When such fraudulent apps activate new communication addresses, they are automatically identified and blocked immediately based on stable characteristics. This eliminates the need for manual analysis of other installation packages with the same framework. Once a victim has installed a fraudulent app, it cannot be used, buying time for dissuasion.
[0097] Step S206: Fraudsters change servers and update the app Once fraudsters discover that a fraudulent app has been blocked and is unusable, they redeploy it on a new server or public cloud, changing superficial characteristics such as the communication server IP address, icon, packet name, and MD5 value. They then repeatedly execute steps S203, S204, and S205. The app is further identified based on stable characteristics within network traffic, and its traffic data is extracted. Each traffic data entry typically contains information such as source address, source port, destination IP address, and destination port. The new destination IP address corresponds to new volatile characteristics of the app (i.e., a new server IP address, etc.). These volatile characteristics are automatically extracted and blocked, with the blocking process continuously iterating and looping.
[0098] In this embodiment, even if fraudsters deploy fraudulent apps on new servers and change superficial features such as the communication server IP address, icon, package name, and MD5 value, they can still identify the fraudulent app on the network through the fingerprint feature database and extract the communication server IP address in real time for blocking.
[0099] Based on the target application identification method provided in the above embodiments, this application also provides specific implementations of the target application identification device. Please refer to the following embodiments.
[0100] like Figure 3 As shown, the target application identification device 300 provided in this application embodiment may include the following modules: a first acquisition module 301, a first extraction module 302, an identification module 303, a first determination module 304, and a blocking module 305.
[0101] The first acquisition module 301 is used to acquire network traffic data of the application. The first extraction module 302 is used to extract stable features of the application based on the application's network traffic data. The stable features are the inherent features of the application based on the development framework. The identification module 303 is used to identify the stable features of the application according to the preset stable feature library and generate identification results. The stable feature library includes the stable features of multiple fraudulent application installation package samples and the labels of each fraudulent application installation package sample. The stable features of different fraudulent application installation package samples are different. The identification results are used to indicate whether the stable features of the application match the stable features of any fraudulent application installation package sample. The first determining module 304 is used to determine the application as the target application if the stable characteristics of the application characterized by the identification result match the stable characteristics of any fraudulent application installation package sample. The blocking module 305 is used to block the target application based on the information of the volatile characteristics of the corresponding tag of the target application in the network traffic data. Different tags have different volatile characteristics.
[0102] The target application identification device of this application embodiment is capable of acquiring network traffic data of the application; extracting stable features of the application based on the network traffic data, wherein stable features are inherent features of the application based on the development framework; identifying the stable features of the application according to a preset stable feature library, generating an identification result, wherein the stable feature library includes stable features of multiple fraudulent application installation package samples and tags of each fraudulent application installation package sample, wherein the stable features of different fraudulent application installation package samples are different, and the identification result is used to indicate whether the stable features of the application match the stable features of any fraudulent application installation package sample; if the identification result indicates that the stable features of the application match the stable features of any fraudulent application installation package sample, the application is determined to be a target application; and blocking the target application based on the information of the volatile features of the tags corresponding to the target application in the network traffic data, wherein the volatile features corresponding to different tags are different. Thus, in this embodiment, since stable features are inherent to the application based on the development framework and are not easily changed, fraudulent applications can be accurately and efficiently identified in network traffic data based on stable features in the stable feature library, overcoming the limitations of static feature identification. Then, targeted blocking is implemented based on the volatile feature information of its corresponding tag. Even if the fraudulent application updates its volatile features, the blocking strategy can still be re-identified and updated through stable features, avoiding the failure of traditional blocking methods based on static features due to feature changes. This achieves continuous and effective prevention and control of fraudulent applications and effectively protects users' property security and personal information rights.
[0103] In some embodiments, the first extraction module 302 described above may specifically include: The parsing unit is used to parse the plaintext header information of the transport layer security protocol handshake phase in the network traffic data of the application and extract the transport layer security protocol metadata. The first extraction unit is used to extract packet sequence features from the network traffic data of the application. The second extraction unit is used to extract the time dimension features from the application's network traffic data; The third extraction unit is used to extract load characteristics from the network traffic data of the application. The aggregation unit is used to aggregate transport layer security protocol metadata, message sequence characteristics, time dimension characteristics, and payload characteristics to obtain stable characteristics of the application.
[0104] In some embodiments, the above-mentioned message sequence features may include packet length sequence, packet length distribution, and packet length direction transition matrix, and the above-mentioned first extraction module 302 may further include: An enhancement unit is used to enhance the packet-length sequence to obtain an enhanced packet-length sequence. The enhancement process includes at least one of discrete Fourier transform, temporal symbolic compression, and state transition probability quantization. The normalization unit is used to normalize the time dimension features to obtain enhanced time dimension features; The aforementioned aggregation unit can be used to aggregate transport layer security protocol metadata, enhanced packet length sequences, packet length distributions, packet length direction transition matrices, enhanced time dimension features, and load features to obtain stable features of the application.
[0105] As one implementation of this application, in order to build a stable feature library, the above-mentioned device 300 may further include: The second acquisition module is used to acquire multiple fraudulent application installation package samples; The second extraction module is used to extract stable features from the installation package samples of each fraudulent application. The second determination module is used to determine the label of each fraudulent application installation package sample based on the stable characteristics of each fraudulent application installation package sample. Different labels of fraudulent application installation package samples correspond to different volatile characteristics. The module is used to build a stable feature library based on multiple fraudulent application installation package samples, the stable features and tags of each fraudulent application installation package sample.
[0106] In some embodiments, the above-mentioned construction module can be used to store multiple fraudulent application installation package samples, stable features and tags of each fraudulent application installation package sample according to a three-layer index architecture of application identifier, software development kit version and network environment type, and build a stable feature library. The aforementioned network traffic data may include application identifiers, software development kit versions, and network environment types. The aforementioned identification module 303 may specifically include: The retrieval unit is used to search a stable feature database based on the application's application identifier, software development kit version, and network environment type to locate target fraudulent application installation package samples that match the application. The identification unit is used to identify the stable features of the application based on the sample of the target fraudulent application installation package, and generate identification results. The identification results are used to indicate whether the stable features of the application match the stable features of the sample of the target fraudulent application installation package.
[0107] In some embodiments, the blocking module 305 may specifically include: The first determining unit is used to determine the tags of fraudulent application installation package samples that match the target application in the stable feature library as the target tags of the target application. The second determining unit is used to determine the volatile features corresponding to the target label as the target volatile features of the target application. The extraction unit is used to extract information about the target's volatile characteristics from network traffic data; The blocking unit is used to block the target application based on information about the target's volatile characteristics.
[0108] In some embodiments, the blocking unit may specifically include: The cleaning and processing subunit is used to clean and process the information of the target's volatile features to obtain the standard information of the target's volatile features. The blocking subunit is used to block the target application based on standard information about the target's volatile characteristics.
[0109] As another implementation of this application, in order to reduce the risk of accidental blocking, the above-mentioned device 500 may further include: The third acquisition module is used to acquire the risk level of at least one user using the target application; The third determination module is used to determine the risk level of the target application based on the risk level of at least one user; The aforementioned blocking module 305 is specifically used to block the target application based on information about the target's volatile characteristics when the risk level of the target application is greater than the preset risk target level.
[0110] Figure 4 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.
[0111] An electronic device may include a processor 401 and a memory 402 storing computer program instructions.
[0112] Specifically, the processor 401 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0113] Memory 402 may include mass storage for data or instructions. For example, and not limitingly, memory 402 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 402 is non-volatile solid-state memory.
[0114] In a particular embodiment, memory 402 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Thus, generally, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of this disclosure.
[0115] The processor 401 reads and executes computer program instructions stored in the memory 402 to implement any of the target application identification methods in the above embodiments.
[0116] In one example, the electronic device may also include a communication interface 403 and a bus 410. For example, Figure 4 As shown, the processor 401, memory 402, and communication interface 403 are connected through bus 410 and complete communication with each other.
[0117] The communication interface 403 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0118] Bus 410 includes hardware, software, or both, that couples components of an electronic device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 410 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.
[0119] The electronic device can execute the target application identification method in the embodiments of this application, thereby achieving the combination Figure 1 and Figure 3 The method and apparatus for identifying the target application are described.
[0120] Furthermore, in conjunction with the target application identification method in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement any of the target application identification methods in the above embodiments.
[0121] This application also provides a computer program product, including a computer program, which, when executed, implements any of the target application identification methods described in the above embodiments.
[0122] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0123] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0124] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0125] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0126] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A method for identifying a target application, characterized in that, include: Obtain network traffic data from the application; Based on the network traffic data of the application, extract the stable features of the application, which are the inherent features of the application based on the development framework; The stable features of the application are identified according to a preset stable feature library, and an identification result is generated. The stable feature library includes the stable features of multiple fraudulent application installation package samples and the tags of each fraudulent application installation package sample. The stable features of different fraudulent application installation package samples are different. The identification result is used to indicate whether the stable features of the application match the stable features of any one of the fraudulent application installation package samples. If the stable characteristics of the application identified by the recognition result match the stable characteristics of any of the fraudulent application installation package samples, the application is determined to be the target application. Based on the information of the volatile characteristics of the target application corresponding to the tag in the network traffic data, the target application is blocked. The volatile characteristics are different for different tags.
2. The method according to claim 1, characterized in that, The step of extracting stable features of the application based on the application's network traffic data includes: The plaintext header information of the transport layer security protocol handshake phase in the network traffic data of the application is parsed to extract the transport layer security protocol metadata; Extract packet sequence features from the network traffic data of the application; Extract the time dimension features from the network traffic data of the application; Extract load characteristics from the network traffic data of the application; The stable characteristics of the application are obtained by aggregating the transport layer security protocol metadata, the message sequence characteristics, the time dimension characteristics, and the payload characteristics.
3. The method according to claim 2, characterized in that, The message sequence features include packet length sequence, packet length distribution, and packet length direction transition matrix. Before aggregating the transport layer security protocol metadata, the message sequence features, the time dimension features, and the payload features to obtain the stable features of the application, the method further includes: The packet-length sequence is enhanced to obtain an enhanced packet-length sequence. The enhancement process includes at least one of discrete Fourier transform, temporal symbolic compression, and state transition probability quantization. The time dimension features are normalized to obtain enhanced time dimension features; The aggregation of the transport layer security protocol metadata, the message sequence characteristics, the time dimension characteristics, and the payload characteristics to obtain the stable characteristics of the application includes: The stable characteristics of the application are obtained by aggregating the transport layer security protocol metadata, the enhanced packet length sequence, the packet length distribution, the packet length direction transition matrix, the enhanced time dimension features, and the load features.
4. The method according to claim 1, characterized in that, Before identifying the stable features of the application based on a preset stable feature library and generating identification results, the method further includes: Obtain sample installation packages of the aforementioned fraudulent applications; Extract stable features from each of the fraudulent application installation package samples; Based on the stable characteristics of each fraudulent application installation package sample, a label is determined for each fraudulent application installation package sample, and the fraudulent application installation package samples with different labels correspond to different volatile characteristics. Based on the multiple fraudulent application installation package samples, the stable features and tags of each fraudulent application installation package sample, the stable feature library is constructed.
5. The method according to claim 4, characterized in that, The stable feature library is constructed based on the multiple fraudulent application installation package samples, the stable features and tags of each fraudulent application installation package sample, including: The stable feature library is constructed by storing the multiple fraudulent application installation package samples, the stable features and tags of each fraudulent application installation package sample according to a three-layer index architecture of application identifier, software development kit version and network environment type. The network traffic data includes application identifiers, software development kit versions, and network environment types. The step of identifying the stability characteristics of the application based on a preset stability feature library and generating identification results includes: Based on the application identifier, software development kit version and network environment type of the application, the target fraudulent application installation package sample that matches the application is searched in the stable feature database; Based on the target fraudulent application installation package sample, the stability characteristics of the application are identified, and an identification result is generated. The identification result is used to indicate whether the stability characteristics of the application match the stability characteristics of the target fraudulent application installation package sample.
6. The method according to claim 1, characterized in that, The step of blocking the target application based on the volatile characteristics of the tag corresponding to the target application in the network traffic data includes: The tags of the fraudulent application installation package samples that match the target application in the stable feature library are determined as the target tags of the target application; The volatile features corresponding to the target label are determined as the target volatile features of the target application. Extract information about the target's volatile characteristics from the network traffic data; Based on the information regarding the target's volatile characteristics, the target application is blocked.
7. The method according to claim 6, characterized in that, The blocking of the target application based on the information of the target's volatile characteristics includes: The information of the target's volatile features is cleaned to obtain standard information of the target's volatile features; The target application is blocked based on the standard information of the target's volatile characteristics.
8. The method according to claim 1, characterized in that, Before blocking the target application based on the volatile characteristics of the target application's corresponding tag in the network traffic data, the method further includes: Obtain the risk level of at least one user using the target application; The risk level of the target application is determined based on the risk level of the at least one user; The step of blocking the target application based on the volatile characteristics of the tag corresponding to the target application in the network traffic data includes: If the risk level of the target application is greater than the preset risk target level, the target application will be blocked based on the information of the target's volatile characteristics.
9. A target application identification device, characterized in that, The device includes: The first acquisition module is used to acquire network traffic data of the application. The first extraction module is used to extract stable features of the application based on the network traffic data of the application. The stable features are the inherent features of the application based on the development framework. The identification module is used to identify the stable features of the application according to a preset stable feature library and generate an identification result. The stable feature library includes the stable features of multiple fraudulent application installation package samples and the tags of each fraudulent application installation package sample. The stable features of different fraudulent application installation package samples are different. The identification result is used to indicate whether the stable features of the application match the stable features of any one of the fraudulent application installation package samples. The first determining module is used to determine the application as the target application if the stable characteristics of the application characterized by the identification result match the stable characteristics of any of the fraudulent application installation package samples. The blocking module is used to block the target application based on the information of the volatile characteristics of the tag corresponding to the target application in the network traffic data. The volatile characteristics are different for different tags.
10. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; the processor, when executing the computer program instructions, implements the target application identification method as described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the method for identifying a target application as described in any one of claims 1-8.
12. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device causes the electronic device to perform the identification method of the target application as described in any one of claims 1-8.