Fraud network resource identification method and device, equipment, medium and product
By identifying candidate domain names in the DPI dataset and constructing a graph neural network for matching and fusion, the problem of low efficiency in identifying fraudulent network resources in traditional methods is solved, achieving fully automated and highly accurate identification of fraudulent resources.
Patent Information
- Application Number
- CN202411494532.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-10-24
AI Technical Summary
Existing technologies have low efficiency in identifying fraudulent online resources, high maintenance and iteration costs for traditional methods, and insufficient intelligence, making it difficult to cope with massive and rapidly changing internet information, resulting in inaccurate identification results and requiring a large amount of manpower.
By acquiring the DPI dataset to be identified, candidate domain names are determined based on strategies such as user time-series behavior and website traffic. A graph neural network is constructed for matching, and graph embedding vectors and content embedding vectors are fused to identify fraudulent resources.
It has achieved fully automated identification of fraudulent online resources, improved identification efficiency and accuracy, broken down the information barriers of variable domain names and IP pools, and effectively responded to new types of hidden fraudulent resources.
Smart Images

Figure CN119561717B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information processing, and in particular to a fraud-related network resource identification method, device, equipment, medium and product. BACKGROUND
[0002] In the field of network anti-fraud, it is necessary to analyze network access data of hundreds of millions of mobile traffic on a daily basis, which puts higher requirements on fraud-related network resource data identification and processing technology. In traditional network anti-fraud business, the main method used is based on blacklists and heuristic rules, etc. These methods have high maintenance iteration costs, and the low level of intelligence makes it difficult to cope with the massive and rapidly changing Internet information, resulting in inaccurate fraud-related network resource data identification results, and only partial automation, which consumes a large amount of manpower and is low in efficiency. SUMMARY
[0003] The present application provides a fraud-related network resource identification method, device, equipment, medium and product to solve the problem of low efficiency of fraud-related network resource identification in the prior art and improve the efficiency of fraud-related network resource identification.
[0004] The present application provides a fraud-related network resource identification method, comprising:
[0005] Obtain a to-be-identified DPI data set, determine a candidate domain name in the domain names in the to-be-identified DPI data set, and the fraud-related possibility of the candidate domain name is higher than that of other domain names in the to-be-identified DPI data set;
[0006] Match the candidate domain name in a pre-set fraud-related domain name library to obtain a diffusion domain name corresponding to the candidate domain name;
[0007] Identify the fraud-related resources of the candidate domain name based on the candidate domain name and the diffusion domain name corresponding to the candidate domain name, and obtain the fraud-related resource identification result of the to-be-identified DPI data set.
[0008] According to the fraud-related network resource identification method provided by the present application, the candidate domain name in the to-be-identified DPI data set is determined, comprising:
[0009] Determine the candidate domain name in the to-be-identified DPI data set based on a first recall strategy and / or a second recall strategy;
[0010] The first recall strategy is determined based on user time sequence behavior, and the second recall strategy is determined based on the access volume and / or access order of the website corresponding to the domain name.
[0011] According to the application, a fraud-related network resource identification method is provided, wherein the candidate domain name is matched in a preset fraud-related domain name library based on the candidate domain name, and a diffusion domain name corresponding to the candidate domain name is obtained.
[0012] A graph is constructed based on the candidate domain name, and a target graph is obtained, wherein nodes in the target graph correspond to the candidate domain name, and edges in the target graph correspond to the association relationship between the candidate domain names.
[0013] The target graph is input into a graph neural network, and a graph embedding vector of each node in the target graph output by the graph neural network is obtained, and the diffusion domain name corresponding to the candidate domain name is obtained by matching the graph embedding vector corresponding to the candidate domain name in the preset fraud-related domain name library.
[0014] According to the application, a fraud-related network resource identification method is provided, wherein a graph is constructed based on the candidate domain name, and a target graph is obtained, comprising:
[0015] The candidate domain name is classified based on the multi-dimensional space statistical features of the candidate domain name, and a plurality of domain name sets are obtained.
[0016] Based on the domain name set, edges in the target graph are generated, and each edge in the target graph connects nodes corresponding to all candidate domain names in the domain name set.
[0017] According to the application, a fraud-related network resource identification method is provided, wherein a graph is constructed based on the candidate domain name, and a target graph is obtained, comprising:
[0018] Based on the DPI data corresponding to the candidate domain name, IP segment aggregation features, user behavior window features, and request path features of the candidate domain name are obtained, and the candidate domain name is connected based on the IP segment aggregation features, the user behavior window features, and the request path features, and edges in the target graph are obtained.
[0019] The IP segment aggregation features reflect the access timestamp interval of the candidate domain names belonging to the same IP segment, the user behavior window features reflect the users corresponding to the candidate domain names belonging to the same time window, and the request path features reflect the access timestamp interval of the candidate domain names sharing the same request path.
[0020] According to the application, a fraud-related network resource identification method is provided, wherein the candidate domain name is matched in a preset fraud-related domain name library based on the candidate domain name, and a diffusion domain name corresponding to the candidate domain name is obtained.
[0021] obtain a content embedding vector corresponding to a target domain name, the content embedding vector reflecting a website content of the target domain name, the target domain name including the candidate domain name and the diffusion domain name corresponding to the candidate domain name;
[0022] fuse the content embedding vector corresponding to the target domain name and the graph embedding vector to obtain a fusion vector;
[0023] determine whether the candidate domain name is a fraud resource based on the fusion vector.
[0024] The application further provides a fraud network resource identification device, comprising:
[0025] a recall module configured to obtain a to-be-identified DPI data set, and determine a candidate domain name in domain names in the to-be-identified DPI data set, the candidate domain name having a higher fraud possibility than other domain names in the to-be-identified DPI data set;
[0026] a diffusion module configured to match the candidate domain name in a preset fraud domain name library to obtain a diffusion domain name corresponding to the candidate domain name;
[0027] an identification module configured to perform fraud resource identification on the candidate domain name based on the candidate domain name and the diffusion domain name corresponding to the candidate domain name, and obtain a fraud resource identification result in the to-be-identified DPI data set.
[0028] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor implementing the fraud network resource identification method when executing the computer program.
[0029] The application further provides a non-transitory computer readable storage medium having a computer program stored thereon, the computer program being executable on a processor to implement the fraud network resource identification method.
[0030] The application further provides a computer program product comprising a computer program, the computer program being executable on a processor to implement the fraud network resource identification method.
[0031] The application provides a fraud-related network resource identification method, device, equipment, medium and product, which comprises the following steps: obtaining a to-be-identified DPI data set, determining a candidate domain name in the domain name in the to-be-identified DPI data set; matching the candidate domain name in a preset fraud-related domain name library to obtain a diffusion domain name corresponding to the candidate domain name; and identifying the candidate domain name and the diffusion domain name corresponding to the candidate domain name as a fraud-related resource to obtain a fraud-related resource identification result in the to-be-identified DPI data set. The application performs preliminary screening on the full amount of DPI data, diffuses the candidate domain name, and takes the candidate domain name and the diffusion domain name with high similarity as the basis for judging whether the candidate domain name is a fraud-related resource, so that the full-automatic identification of the fraud-related network resource can be realized under the premise of ensuring the accuracy of the fraud-related resource identification result, and the efficiency of the fraud-related network resource identification is improved. BRIEF DESCRIPTION OF DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative effort.
[0033] Figure 1 FIG. 1 is a flowchart of the fraud-related network resource identification method provided by the application.
[0034] Figure 2 FIG. 2 is a schematic diagram of determining a candidate domain name in the fraud-related network resource identification method provided by the application.
[0035] Figure 3 FIG. 3 is a schematic diagram of constructing a target graph in the fraud-related network resource identification method provided by the application. Figure 1 .
[0036] Figure 4 FIG. 4 is a schematic diagram of constructing a target graph in the fraud-related network resource identification method provided by the application. Figure 2 .
[0037] Figure 5 FIG. 5 is a schematic diagram of fusing multi-modal data in the fraud-related network resource identification method provided by the application.
[0038] Figure 6 FIG. 6 is a structural schematic diagram of the fraud-related network resource identification device provided by the application.
[0039] Figure 7 FIG. 7 is a structural schematic diagram of the electronic equipment provided by the application. DETAILED DESCRIPTION
[0040] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the present application.
[0041] The present application will be described below in conjunction with Figures 1-5 The present application provides a method for identifying fraudulent network resources. As shown in the figure, the method for identifying fraudulent network resources comprises the following steps: Figure 1
[0042] S110, obtaining a to-be-identified DPI data set, determining a candidate domain name in the domain names in the to-be-identified DPI data set, the fraudulent possibility of the candidate domain name being higher than that of other domain names in the to-be-identified DPI data set;
[0043] S120, matching the candidate domain name in a pre-set fraudulent domain name library to obtain a diffusion domain name corresponding to the candidate domain name;
[0044] S130, identifying the fraudulent resources of the candidate domain name based on the candidate domain name and the diffusion domain name corresponding to the candidate domain name to obtain the identification result of the fraudulent resources in the to-be-identified DPI data set.
[0045] The DPI data set is data obtained by using the DPI (Deep Packet Inspection, deep packet inspection) technology to analyze network messages. The DPI data set comprises a plurality of data, and each piece of data comprises a userID, a timestamp, a host, a route, an IP and an area code. The to-be-identified DPI data set can be obtained by analyzing network messages in a period of time before the current time. After obtaining the to-be-identified DPI data set, the domain names possibly involved in fraud are extracted as candidate domain names.
[0046] The candidate domain name can be determined in the domain names in the to-be-identified DPI data set by a database matching method, that is, a domain name library with high fraudulent possibility is pre-set, and the domain names in the pre-set domain name library with high fraudulent possibility are extracted from the to-be-identified DPI data set as candidate domain names. However, a fraud gang can use highly similar website frameworks and interfaces to build information barriers through variable domain names and IP pools. In the method provided by the present application, in order to prevent omission of the identification of fraudulent resources and improve the accuracy of the candidate domain names, the candidate domain names are extracted based on user behavior and website behavior. Specifically, the candidate domain names are determined in the domain names in the to-be-identified DPI data set, comprising:
[0047] Candidate domain names are determined in the DPI dataset to be identified based on the first recall strategy and / or the second recall strategy;
[0048] The first recall strategy is determined based on user temporal behavior, while the second recall strategy is determined based on the website traffic and / or access order corresponding to the domain name.
[0049] Since the fraud risks in DPI (Data Points Intake) internet browsing logs are often highly correlated with factors such as visits to risky websites, online banking, and overseas access, the method provided in this invention uses multi-granularity internet browsing log sequence analysis to deeply mine slice-level samples of the original DPI data. This extracts information reflecting users' long-term behavioral habits, access to niche customer service systems, access to overseas risky IP addresses, abnormal access fluctuations, abnormal download behavior, and suspected payment behavior. From this, a fraud-related resource recall strategy is constructed. Based on the recall strategy, candidate domains with a high probability of fraud are recalled from the DPI dataset to be identified for further detailed judgment.
[0050] Specifically, the method provided by this invention includes a recall strategy for retrieving candidate domain names from the DPI dataset to be identified, which may include a first recall strategy and / or a second recall strategy. The first recall strategy is determined based on user temporal behavior, and the second recall strategy is determined based on the number of visits to the website corresponding to the domain name and / or the order of visits. Figure 2 As shown, the first recall strategy based on user behavior summarizes the temporal behavior of alerted users, identifies fraudulent patterns, and extracts suspicious domain names within a specified time range. The second recall strategy based on website behavior directly extends the identification dimension to the website, performing indicator calculations and temporal context detection to determine suspicious domain names. When determining candidate domain names based on the first and / or second recall strategies, after obtaining the DPI dataset to be identified, the domain names in the DPI dataset are matched with the suspicious domain names based on the first and / or second recall strategies to extract candidate domain names.
[0051] Specifically, since fraudulent domain names are often associated with fake customer service chats between criminals and victims, a set of... A collection of domains related to chat software:
[0052] {'%antchats.im%','%ya.cn%','%smallchat.chat%','%isweetalk.com%','%zhifeiji.chat%','%ourchat.com.cn%'}.
[0053] Simultaneously define Define a time window function for the current time. For any DPI record , Returns a boolean value representing the record. Is the timestamp in time? Previous Within a time period (can be set) =5min).
[0054] Thus, the first recall strategy can be expressed as:
[0055] .
[0056] That is, the first recall strategy is to select the time elapsed since the current time. Overseas domains related to chat software that were visited within the specified time period were selected as candidate domains.
[0057] For the second recall strategy, it can be built based on the number of visits and / or the order of visits. Building based on the number of visits involves setting thresholds for indicators such as the number of visits and the number of users visiting the website based on statistical knowledge of existing cases to form rules (e.g., domains with fewer than 100 users visiting in the past 168 hours). Building based on the order of visits involves matching the website according to the fingerprint paths commonly used in the same fraud scenario, and performing pattern matching of "A first, then B, then C" (e.g., domains that appear one after another within a 35-minute window, such as weixin110 and meiqia).
[0058] Based on the first and second recall strategies, after extracting candidate domain names from the DPI data to be identified, further processing is performed to determine whether the candidate domain names are fraudulent network resources.
[0059] Currently, scammers no longer consider phone calls and text messages as essential communication methods, instead relying more on various general or customized apps and websites. However, carrier DPI data cannot reveal the encrypted communication content within these applications. This results in severe sparsity in graphs constructed using traditional methods, which cannot effectively depict the interaction between fraudulent websites and users. Existing technical solutions treat each fraudulent resource (such as a website address or app) independently and perform fraud detection based on this. In reality, to carry out telecommunications fraud, a large amount of interactive information is generated between fraudulent resources and victims, primarily through the internet (including accessing web pages and using mobile applications). However, this communication relationship information is difficult to effectively utilize in previous detection methods based on structured data (tabular data). Figure 3 As shown, the method provided by this invention considers large-scale mapping of the full DPI data to mine the connections between fraudulent resources, performs association diffusion on individual candidate domains, extracts known fraudulent resources with a high degree of association, and improves the accuracy of the fraudulent resource identification results of candidate domains.
[0060] Specifically, the candidate domain name is matched in the preset fraud-related domain name library based on the candidate domain name to obtain a diffusion domain name corresponding to the candidate domain name, including:
[0061] A target graph is obtained based on the candidate domain name, nodes in the target graph correspond to the candidate domain name, and edges in the target graph correspond to an association relationship between the candidate domain names;
[0062] The target graph is input into a graph neural network, and a graph embedding vector of each node in the target graph output by the graph neural network is obtained, and the diffusion domain name corresponding to the candidate domain name is obtained based on the matching of the graph embedding vector corresponding to the candidate domain name in the preset fraud-related domain name library.
[0063] In a possible implementation, for full-DPI data, edges are generated using IP segments, user behaviors, and http requests, network resource domain names are vertices, and large-scale graph construction is performed on three properties of IP segment aggregation, interface multiplexing, and request specificity in DPI user online behaviors, as shown in Figure 4 By exploring the specificity of the website interface and the multiplexing of the user behavior interface, more potential risk domain names and websites can be explored. A target graph is obtained based on the candidate domain name, including:
[0064] Based on the DPI data corresponding to the candidate domain name, IP segment aggregation features, user behavior window features, and request path features of the candidate domain name are obtained, the candidate domain names are connected based on the IP segment aggregation features, the user behavior window features, and the request path features, and edges in the target graph are obtained;
[0065] The IP segment aggregation features reflect the access timestamp interval of the candidate domain names belonging to the same IP segment, the user behavior window features reflect the users corresponding to the candidate domain names belonging to the same time window, and the request path features reflect the access timestamp interval of the candidate domain names sharing the same request path.
[0066] In this implementation, the target graph is a normal graph, and one edge connects at most two vertices. In this implementation, the formulaic expression for constructing the target graph is as follows:
[0067] Input: DPI data set corresponding to the candidate domain name , containing user_id, timestamp, host, uri, IP, etc.
[0068] Output: graph , where V is a node set (i.e., a domain name set), and E is an edge set.
[0069] Parameters: T: time window size (e.g., 5 minutes); D: continuous observation days (e.g., 3 days).
[0070] Initialization:
[0071]
[0072]
[0073] Preprocessing:
[0074] For each record , add the corresponding each host to the node set V.
[0075] ① Edge construction based on IP aggregation:
[0076] Group all records by IP prefix, i.e. . For records within each IP segment , if their timestamps are within D days, and if two different domain names and appear in the same IP segment, add an edge to the edge set E (if the edge does not exist yet).
[0077] ② Edge construction based on user behavior window:
[0078] For each user user_id, sort and group records by timestamp, obtaining . For each user's time window (last T minutes), if two different domain names and are accessed by the same user within the same time window, add an edge to the edge set E (if the edge does not exist yet).
[0079] ③ Edge construction based on request path:
[0080] For all specific request paths (artificially determined specific phishing request paths), find the record set containing all uris of the path, obtaining . For records within each specific request path , if there are two different domain names and , whose timestamps are within D days, and uris have the same request path, add an edge to the edge set E (if the edge does not exist yet).
[0081] Return: target graph .
[0082] Because a typical graph's edge connects only two nodes, it struggles to represent unpaired relationships. Criminal gangs, using highly similar website frameworks and interfaces, construct information barriers through variable domain names and IP pools. Addressing the issue of insufficient global high-order relationships caused by these fraudulent resource groups, this invention provides another implementation that employs a heterogeneous hypergraph network embedding method, where the target graph is a hypergraph. A hypergraph is a special graph structure where an edge can contain any number of nodes, making it an efficient way to uniformly model various unpaired relationships. In a simple graph, each edge is associated with two vertices, meaning the degree of each edge is limited to 2. A hypergraph, however, allows the degree of each edge to be any non-negative integer. To represent the global associations between the recalled candidate domains, this implementation uses a heterogeneous hypergraph network embedding method. Based on the multidimensional domain name space features, a hypergraph is formed. ( (a diagonal matrix of the weights of each hyperedge), through a hypergraph structure By grouping domains with high-level relationships into a super-edge area, a high-level global view of the entire network can be provided, including various relationships such as IP addresses and user behavior. This breaks down the information barriers built by variable domain names and IP pools, making it difficult for fraudsters to hide all traces of their actions and effectively countering new types of hidden fraudulent resources.
[0083] In the implementation based on hypergraph network embedding, a graph is constructed based on candidate domain names to obtain the target graph, including:
[0084] Candidate domain names are classified based on their multidimensional spatial statistical characteristics, resulting in multiple domain name sets.
[0085] Edges in the target graph are generated based on the domain name set, and each edge in the target graph connects to the nodes corresponding to all candidate domain names in the domain name set.
[0086] In hypergraph algorithms, the quality of hyperedge construction directly affects the overall recognition performance of the model. Considering that the spatial statistical features of domain names suspected of fraud are mostly numerical and categorical features, one implementation of the method provided in this invention considers extracting multidimensional spatial statistical features of the domain name (including but not limited to the domain name length, the proportion of numbers in the domain name, the length of the longest consecutive number in the domain name, the offset value of the longest consecutive number sequence, the number of WHOIS registrars, the number of registered countries, the number of name servers, the number of registrants, the registration year, registration month, registration day, registration time, registration minute, registration second, expiration year, expiration month, expiration day, expiration hour, expiration minute, expiration second, last update year, last update...). Based on the features of the new moon, the last update date, the last update time, the last update minute, the last update second, the number of visitors in the last 1 day, the number of visitors in the last 7 days, the number of visitors in the last 30 days, the number of DPI records in the last 1 day, the number of DPI records in the last 7 days, and the number of DPI records in the last 30 days, an adaptive hypergraph is constructed by training a tree-based model (including but not limited to decision trees, random forests, GBDT, etc.). After obtaining the tree model with optimal parameters, all candidate domain names from the previous stage are pre-classified (divided into C categories according to the prediction results of the second-to-last leaf nodes of the tree model). Then, a set of hyperedges E is formed based on the classification results. E contains C hyperedges, and each hyperedge corresponds to a set of candidate domain names for a category.
[0087] After obtaining the target graph, a graph neural network is used to obtain the embedding vector representation of the domain name nodes. If the target graph is a simple graph, a graph neural network for simple graphs is used; if the target graph is a hypergraph, a hypergraph neural network is used. The graph neural network includes multiple layers of graph convolution operations to perform convolution transformations on the encoded features of domain name characters, further enhancing the domain name features and fully representing the global relationships between domain names. The hypergraph convolution operation formula is as follows:
[0088]
[0089] in, For the first Embedding vectors of layered neural networks and Let represent the diagonal matrices of the vertex degree matrix and the hyperedge degree matrix of the hypergraph, respectively. It is a non-linear activation function. The parameters to be learned during training are updated using the cross-entropy loss function through backpropagation.
[0090] After obtaining the graph embedding vector of the candidate domain name through graph neural network mapping, the similarity between the graph embedding vector of the candidate domain name and the fraudulent domain name vector in the historical black database is calculated, and the top N most similar fraudulent domain names are taken as the diffusion domain names corresponding to the candidate domain name.
[0091] The vector representation of the domain name related to fraud in the historical black library is also obtained through the graph deep network mapping, and the graph neural network can be trained by using the sample domain name and the similarity label between the sample domain names, so that the graph neural network can output similar vector representations for similar domain names.
[0092] Based on the candidate domain name and the diffusion domain name corresponding to the candidate domain name, the candidate domain name is identified as a resource related to fraud, including:
[0093] Obtaining a content embedding vector corresponding to the target domain name, the content embedding vector reflecting the content of the target domain name, the target domain name including the candidate domain name and the diffusion domain name corresponding to the candidate domain name;
[0094] Fusing the content embedding vector and the graph embedding vector corresponding to the target domain name to obtain a fusion vector;
[0095] Determining whether the candidate domain name is a resource related to fraud based on the fusion vector.
[0096] The content embedding vector of the target domain name can be an embedding vector of content including two modalities (text and picture). Specifically, the snapshot screenshot of the website and the HTML content of the website can be obtained through a snapshot representation model (including but not limited to ViT) and a text representation model (including but not limited to BERT) respectively to obtain the corresponding embedding vector as the content embedding vector. The content embedding vector and the graph embedding vector corresponding to the target domain name are denoted as , respectively. The vector length of each modality is denoted as , and the vector is obtained by concatenating , where c is the sum of Figure 5 . The content embedding vector and the graph embedding vector corresponding to the target domain name are fused to obtain a fusion vector; the process of determining whether the candidate domain name is a resource related to fraud based on the fusion vector can be realized through a trained neural network, and the training data of the neural network includes sample target domain names and labels of whether they are resources related to fraud. As shown in , through the neural network, based on the concatenated vector Z, three different are obtained. , and then multiplied by to obtain , and then concatenated to obtain the overall modal fusion vector . The final result is obtained by performing a softmax layer classification, that is, the judgment result of whether the candidate domain name is a resource related to fraud is obtained.
[0097] Through modal fusion, not only the mutual dependence between nodes can be represented, but also the multi-dimensional space statistical characteristics of nodes or edges can be fused, and the topological structure information and attribute information of the resource related to fraud can be represented at the same time, thereby improving the expression ability of the model and the accuracy of the identification result of the resource related to fraud.
[0098] The fraud-related network resource identification device provided by the present application is described below, and the fraud-related network resource identification device described below can be referred to each other corresponding to the fraud-related network resource identification method described above. As shown in Figure 6 The fraud-related network resource identification device provided by the present application includes:
[0099] The recall module 610 is configured to obtain a to-be-identified DPI data set, determine a candidate domain name in the domain name in the to-be-identified DPI data set, and the fraud-related possibility of the candidate domain name is higher than that of other domain names in the to-be-identified DPI data set.
[0100] The diffusion module 620 is configured to match the candidate domain name in a preset fraud-related domain name library to obtain a diffusion domain name corresponding to the candidate domain name.
[0101] The identification module 630 is configured to perform fraud-related resource identification on the candidate domain name based on the candidate domain name and the diffusion domain name corresponding to the candidate domain name, and obtain a fraud-related resource identification result in the to-be-identified DPI data set.
[0102] Figure 7 An example of an entity structure diagram of an electronic device is shown in Figure 7 The electronic device can include a processor 710, a communications interface 720, a memory 730, and a communications bus 740, wherein the processor 710, the communications interface 720, and the memory 730 can communicate with each other through the communications bus 740. The processor 710 can call the logical instructions in the memory 730 to execute the fraud-related network resource identification method, which includes: obtaining a to-be-identified DPI data set, determining a candidate domain name in the domain name in the to-be-identified DPI data set, and the fraud-related possibility of the candidate domain name is higher than that of other domain names in the to-be-identified DPI data set; matching the candidate domain name in a preset fraud-related domain name library to obtain a diffusion domain name corresponding to the candidate domain name; performing fraud-related resource identification on the candidate domain name based on the candidate domain name and the diffusion domain name corresponding to the candidate domain name, and obtaining a fraud-related resource identification result in the to-be-identified DPI data set.
[0103] Further, the logic instructions in the memory 730 described above can be implemented in the form of software functional units and sold or used as standalone products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or partially contribute to the prior art, or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0104] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the fraud-related network resource identification method provided by the above-mentioned methods. The method comprises: obtaining a to-be-identified DPI data set, determining a candidate domain name in the domain name in the to-be-identified DPI data set, and the fraud-related possibility of the candidate domain name is higher than that of other domain names in the to-be-identified DPI data set; matching the candidate domain name in a preset fraud-related domain name library to obtain the diffusion domain name corresponding to the candidate domain name; and performing fraud-related resource identification on the candidate domain name based on the candidate domain name and the diffusion domain name corresponding to the candidate domain name, to obtain the fraud-related resource identification result in the to-be-identified DPI data set.
[0105] In another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the fraud-related network resource identification method provided by the above-mentioned methods. The method comprises: obtaining a to-be-identified DPI data set, determining a candidate domain name in the domain name in the to-be-identified DPI data set, and the fraud-related possibility of the candidate domain name is higher than that of other domain names in the to-be-identified DPI data set; matching the candidate domain name in a preset fraud-related domain name library to obtain the diffusion domain name corresponding to the candidate domain name; and performing fraud-related resource identification on the candidate domain name based on the candidate domain name and the diffusion domain name corresponding to the candidate domain name, to obtain the fraud-related resource identification result in the to-be-identified DPI data set.
[0106] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0107] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0108] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying fraudulent online resources, characterized in that, include: Obtain the DPI dataset to be identified, and determine candidate domains from the domains in the DPI dataset, wherein the candidate domains are more likely to be involved in fraud than other domains in the DPI dataset. Based on the candidate domain name, a matching process is performed in a preset database of fraudulent domain names to obtain the diffusion domain name corresponding to the candidate domain name; Based on the candidate domain name and the diffusion domain name corresponding to the candidate domain name, the candidate domain name is used to identify fraudulent resources, and the fraudulent resource identification result in the DPI dataset to be identified is obtained. The process of matching the candidate domain names in a preset database of fraudulent domain names to obtain the corresponding dissemination domain names includes: A target graph is obtained by constructing a graph based on the candidate domain names, where the nodes in the target graph correspond to the candidate domain names and the edges in the target graph correspond to the associations between the candidate domain names. The target graph is input into a graph neural network to obtain the graph embedding vector of each node in the target graph output by the graph neural network. Based on the graph embedding vector corresponding to the candidate domain name, a matching is performed in a preset fraud-related domain name database to obtain the diffusion domain name corresponding to the candidate domain name.
2. The method for identifying fraudulent online resources according to claim 1, characterized in that, The step of determining candidate domain names from the domain names in the DPI dataset to be identified includes: Candidate domain names are determined in the dataset of DPIs to be identified based on the first recall strategy and / or the second recall strategy. The first recall strategy is determined based on user temporal behavior, while the second recall strategy is determined based on the website traffic and / or access order corresponding to the domain name.
3. The method for identifying fraudulent online resources according to claim 1, characterized in that, The process of constructing a graph based on the candidate domain names to obtain the target graph includes: Based on the multidimensional spatial statistical characteristics of the candidate domain names, the candidate domain names are classified to obtain multiple domain name sets; Edges in the target graph are generated based on the domain name set, and each edge in the target graph connects to a node corresponding to all the candidate domain names in the domain name set.
4. The method for identifying fraudulent online resources according to claim 1, characterized in that, The process of constructing a graph based on the candidate domain names to obtain the target graph includes: Based on the DPI data corresponding to the candidate domain names, the IP segment clustering characteristics, user behavior window characteristics, and request path characteristics of the candidate domain names are obtained. The candidate domain names are then connected based on the IP segment clustering characteristics, user behavior window characteristics, and request path characteristics to obtain the edges in the target graph. The IP segment clustering feature reflects the access timestamp interval of the candidate domains belonging to the same IP segment, the user behavior window feature reflects the users corresponding to the candidate domains belonging to the same time window, and the request path feature reflects the access timestamp interval of the candidate domains sharing the same request path.
5. The method for identifying fraudulent online resources according to claim 1, characterized in that, The step of identifying fraudulent resources based on the candidate domains and the corresponding dissemination domains includes: Obtain the content embedding vector corresponding to the target domain name, the content embedding vector reflecting the URL content of the target domain name, the target domain name including the candidate domain name and the diffusion domain name corresponding to the candidate domain name; The content embedding vector and the graph embedding vector corresponding to the target domain name are fused to obtain a fused vector; The candidate domain name is determined as a fraudulent resource based on the fusion vector.
6. A device for identifying fraudulent online resources, characterized in that, include: The recall module is used to obtain the DPI dataset to be identified, and to determine candidate domains among the domains in the DPI dataset, wherein the candidate domains are more likely to be involved in fraud than other domains in the DPI dataset. The diffusion module is used to match the candidate domain name in a preset database of fraudulent domain names to obtain the diffusion domain name corresponding to the candidate domain name; The identification module is used to identify fraudulent resources in the candidate domain name based on the candidate domain name and the diffusion domain name corresponding to the candidate domain name, and to obtain the fraudulent resource identification result in the DPI dataset to be identified. The process of matching the candidate domain names in a preset database of fraudulent domain names to obtain the corresponding dissemination domain names includes: A target graph is obtained by constructing a graph based on the candidate domain names, where the nodes in the target graph correspond to the candidate domain names and the edges in the target graph correspond to the associations between the candidate domain names. The target graph is input into a graph neural network to obtain the graph embedding vector of each node in the target graph output by the graph neural network. Based on the graph embedding vector corresponding to the candidate domain name, a matching is performed in a preset fraud-related domain name database to obtain the diffusion domain name corresponding to the candidate domain name.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the fraud-related network resource identification method as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the fraud-related network resource identification method as described in any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the fraud-related network resource identification method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Illegal application software identification method and device, medium and electronic equipment
CN113890866A
Campus network fraud early warning method and system based on DPI, and readable medium
CN118523969A