A fraud information detection method fusing propagation relationships
By filtering and calculating the number of similar text information and fraud weight values, the problem of low efficiency in fraud information detection in existing technologies has been solved, and efficient fraud information detection has been achieved.
Patent Information
- Application Number
- CN202211600947.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-13
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2042-12-13
AI Technical Summary
Existing information processing technologies are unable to efficiently detect large amounts of fraudulent information, resulting in low processing efficiency.
By acquiring a database of fraudulent accounts and a database of legitimate accounts, information groups that do not belong to either database are filtered out. The number of similar text messages from the sending accounts in the information groups is calculated, and a fraud weight value is calculated to identify fraudulent information.
It improves the processing efficiency of fraud information detection, enabling efficient detection of large amounts of fraud information.
Smart Images

Figure CN116127964B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computers, and more specifically, to a method for detecting fraudulent information based on integrated propagation relationships. Background Technology
[0002] In recent years, with the rapid development of information technology, the deployment and use of the internet have been rapidly and widely popularized. A vast amount of information from different fields, regions, and time zones is disseminated extensively via the internet immediately after its creation. Simultaneously, the in-depth development of mobile internet and telecommunications networks, along with the widespread adoption of mobile handheld communication devices, has further amplified this phenomenon. However, information on the internet is generated without any verification, therefore its authenticity cannot be guaranteed. A large amount of exaggerated, false, or even fabricated information is mixed with true information, often making it difficult for people to distinguish between them. The number of false information generated in a short period is in the hundreds of millions, and traditional information processing technologies and analytical methods are unable to handle the processing and computation of this massive data volume. Summary of the Invention
[0003] This invention provides a method for detecting fraudulent information based on integrated propagation relationships, thereby addressing the technical problem of low processing efficiency when detecting large amounts of fraudulent information.
[0004] According to one aspect of the present invention, a method for detecting fraudulent information based on fusion propagation relationships is provided, comprising: acquiring a first information group, a fraudulent account database, and a normal account database, wherein each piece of information in the first information group includes text information and a sending account; determining a second information group from the first information group based on the fraudulent account database and the normal account database, wherein the sending account of each piece of information in the second information group does not exist in either the fraudulent account database or the normal account database; obtaining a plurality of target information groups based on the second information group, wherein the number of similar text information between a first sending account and a second sending account in each target information group is greater than a first threshold; calculating a fraud weight value for each target information group; and determining each piece of text information in the target information group as fraudulent information when the fraud weight value of the target information group is greater than a second threshold.
[0005] In this embodiment of the invention, a method is employed to obtain a first information group, a fraudulent account database, and a normal account database, wherein each piece of information in the first information group includes text information and a sending account; a second information group is determined from the first information group based on the fraudulent account database and the normal account database, wherein the sending account of each piece of information in the second information group does not exist in either the fraudulent account database or the normal account database; multiple target information groups are obtained based on the second information group, wherein the number of similar text information between the first and second sending accounts in each target information group is greater than a first threshold; a fraud weight value is calculated for each target information group; and if the fraud weight value of the target information group is greater than a second threshold, each piece of text information in the target information group is identified as fraudulent information. Since the above method involves filtering based on the fraudulent account database and the normal account database in the first step, merging multiple target information groups based on the number of similar text information in the second step, calculating the fraud weight value for each target information group in the third step, and finally identifying all text information in the target information group with a fraud weight value greater than a second threshold as fraudulent information, this method achieves the goal of improving the processing efficiency of detecting fraudulent information and solves the technical problem of low processing efficiency when detecting a large amount of fraudulent information. Attached Figure Description
[0006] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0007] Figure 1 This is a flowchart of an optional method for detecting fraudulent information based on the fusion propagation relationship according to an embodiment of the present invention;
[0008] Figure 2 This is an overall flowchart of an optional method for detecting fraudulent information based on the fusion propagation relationship according to an embodiment of the present invention;
[0009] Figure 3 This is a propagation network diagram of an optional method for detecting fraudulent information based on the fusion propagation relationship according to an embodiment of the present invention. Detailed Implementation
[0010] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0011] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0012] According to a first aspect of the present invention, a method for detecting fraudulent information based on fusion propagation relationships is provided, optionally, as follows: Figure 1 As shown, the above method includes:
[0013] S102, Obtain the first information group, the fraudulent account database, and the normal account database, wherein each piece of information in the first information group includes text information and the sending account;
[0014] S104, determine the second information group from the first information group based on the fraud account database and the normal account database, wherein the sending account of each message in the second information group does not exist in either the fraud account database or the normal account database;
[0015] S106, multiple target information groups are obtained based on the second information group, wherein the number of similar text information between the first sending account and the second sending account in each target information group is greater than the first threshold.
[0016] S108, calculate the fraud weight value for each target information group;
[0017] S110, if the fraud weight value of the target information group is greater than the second threshold, each text information in the target information group is identified as fraud information.
[0018] Optionally, in this embodiment, the first information group comprises all information to be detected, and each piece of information includes text information and the sending account. The fraud account database contains all confirmed fraud accounts, and the legitimate account database contains all confirmed legitimate accounts.
[0019] Optionally, in this embodiment, the first information group is first filtered to select a second information group whose sending account is neither in the normal account database nor the fraudulent account database. Text information from sending accounts in the fraudulent account database is identified as fraudulent information, while text information from sending accounts in the normal account database is identified as normal information. The second information group is then merged based on the number of similar texts. For example, if the first target information group includes sending account 1 and sending account 2, and 10 text messages from sending account 1 have a text similarity exceeding 0.8 with 10 being greater than a first threshold, and so on, multiple target information groups are obtained. Finally, a fraud weight value is calculated for each target information group, and all text information in the target information group with a fraud weight value greater than a second threshold is identified as fraudulent information.
[0020] Optionally, in this embodiment, the first step is to filter based on the fraud account database and the normal account database; the second step is to merge multiple target information groups based on the number of similar text information; the third step is to calculate the fraud weight value of each target information group; and finally, all text information in the target information group with a fraud weight value greater than a second threshold is identified as fraud information. This achieves the goal of improving the processing efficiency of detecting fraud information and solves the technical problem of low processing efficiency when detecting a large amount of fraud information.
[0021] As an optional example, before obtaining multiple target information groups based on the second information group, the above method further includes:
[0022] Calculate the text similarity between the first text information and the second text information, where the first text information is the text information in the second information group, the second text information is the text information in the second information group excluding the first text information, and the first text information and the second text information are text information from different sending accounts;
[0023] If the text similarity is greater than the third threshold, the first text information and the second text information are determined to be similar text information.
[0024] Optionally, in this embodiment, the text similarity between each first text information in the second information group and the second text information of different sending accounts of the first text information is calculated. For example, the text similarity between text information 1 of sending account 1 and text information 1 of sending account 2 is calculated to be 0.8, where 0.8 is greater than the third threshold, and the text information 1 of sending account 1 and text information 1 of sending account 2 are determined to be similar text information.
[0025] As an optional example, calculating the text similarity between the first text information and the second text information includes:
[0026] A first similarity value is determined based on the webpage address in the first text information and the second text information;
[0027] The second similarity value is determined based on the account information in the first and second text information;
[0028] The third similarity value is determined based on the semantics of the first and second text information;
[0029] The sum of the first similarity value, the second similarity value, and the third similarity value is determined as the text similarity score.
[0030] Optionally, in this embodiment, the text content in each text information may include web addresses, accounts, etc. A first similarity value is determined based on the web addresses in the text content of the first and second text information, a second similarity value is determined based on the accounts in the text content of the first and second text information, a third similarity value is determined based on the semantics of the text content of the first and second text information, and finally the sum of the first similarity value, the second similarity value, and the third similarity value is determined as the text similarity between the first and second text information.
[0031] As an optional example, determining the first similarity value based on the webpage address in the first text information and the second text information includes:
[0032] Obtain the first webpage address from the first text information and the second webpage address from the second text information;
[0033] If the first webpage address and the second webpage address are the same, the first similarity value is determined as the first value.
[0034] If the first webpage address and the second webpage address are different, the first similarity value is determined to be the second value.
[0035] Optionally, in this embodiment, if both the first text information and the second text information contain web page addresses, the first web page address and the second web page address of the first text information and the second text information are obtained. If the two web page addresses are the same, a first similarity value is determined as a first value, which can be 1 or greater than a third threshold. That is, when the two web page addresses are the same, the first text information and the second text information are determined to be similar text information. If the two web page addresses are not the same, or neither text information contains a web page address, or one text information does not contain a web page address, the first similarity value is determined as a second value, which can be 0.
[0036] As an optional example, determining the second similarity value based on the account in the first text information and the second text information includes:
[0037] Retrieve the first account from the first text information and the second account from the second text information;
[0038] If the first account and the second account are the same, the second similarity value is determined as the first value;
[0039] If the first account and the second account are different, the second similarity value is determined to be the first value.
[0040] Optionally, in this embodiment, if accounts exist in both the first and second text information, the first and second accounts of the first and second text information are obtained. If the two accounts are the same, a second similarity value is determined to be a first value, which can be 1 or greater than a third threshold. That is, when the two accounts are the same, the first and second text information are determined to be similar text information. If the two accounts are different, or neither text information contains an account, or one text information does not contain an account, a second similarity value is determined to be a second value, which can be 0.
[0041] As an optional example, determining the third similarity value based on the semantics of the first and second text information includes:
[0042] The first text information is segmented into words to obtain the first word group; the second text information is segmented into words to obtain the second word group.
[0043] The third similarity value is calculated based on the first word group, the second word group, and the word embedding model.
[0044] Optionally, in this embodiment, the first text information and the second text information are segmented into word groups to obtain a first word group and a second word group. For example, if the first text information is "Dear Mr. / Ms. Zhang, your credit card points have reached 290068 and will be cleared on the 17th. Please log in to the official website to redeem gifts www.zgdhcc.com [Bank Name]", segmenting it yields the first word group "Dear, Mr. / Ms. Zhang, your, credit card, points, log in to the official website, redeem gifts...". A third similarity value is then calculated based on the first word group, the second word group, and the word embedding model. The word embedding model can be the word2vec algorithm, which is a predictive model that can perform highly efficient word embedding learning.
[0045] As an optional example, calculating the fraud weight value for each target information group includes:
[0046] Each target information group is designated as the current target information group, and the following operations are performed on the current target information group:
[0047] Determine the number of sending accounts in the current target message group;
[0048] The fraud weight value of the current target information group is calculated based on the quantity and random walk model.
[0049] Optionally, in this embodiment, the second information group is converted into multiple network propagation graphs, with the sending account as the central node and the text information of the sending account as the edge nodes. After merging, a target information group becomes a propagation network graph. The central node includes all sending accounts in the target information group, and the edge nodes include all text information in the target information group. The fraud weight value of each target information group is calculated based on the number of sending accounts in each target information group and the random walk pattern. The random walk pattern can be the PageRank algorithm. The PageRank algorithm defines a random walk model on a directed graph, i.e., a first-order Markov chain, describing the behavior of a random walker randomly visiting various nodes along the directed graph. The fraud weight value of the central node is dynamically calculated through the propagation relationship between different nodes.
[0050] To illustrate with an example, this application relates to a method for detecting fraudulent information that integrates propagation relationships. It introduces text vector techniques to accelerate text analysis and improve accuracy, while utilizing text similarity to compensate for the sparsity of the original propagation network structure. Based on this, an improved PageRank algorithm is used to calculate user node weights and mine representative fraudulent information text samples.
[0051] This method has the following advantages:
[0052] 1. Applying natural language processing technologies such as word2vec, text edit distance, NLPIR (Chinese word segmentation system), and Viterbi algorithm, we statistically analyze streaming text information data from a semantic perspective and dynamically calculate text similarity. This accelerates and improves text recognition operations and enhances similarity calculation, making it easier to detect the text prototypes of fraudulent information.
[0053] 2. Application of graph computing related technologies; fraud subgraph detection, by setting a text similarity threshold to merge highly similar text information and sending account, detects fraud subgraphs at different granularities, solves the problem of sparse information dissemination, and constructs a more compact dissemination graph;
[0054] 3. The algorithm has multiple adjustable parameters, the system has feedback operation, and can be set according to requirements. The algorithm has low internal coupling and good portability.
[0055] This method includes the following process, the specific flow of which is as follows: Figure 2 As shown:
[0056] 1. Streaming text data cleaning, keyword extraction, and structured representation. After streaming data is input into the lower-level module, it first undergoes preliminary text processing to generate structured data with consistent attributes. This process includes a series of data operations such as data cleaning, data completion, and blacklist / whitelist filtering. The purpose is to filter out a large amount of text data that does not need to be analyzed and reduce the computational load of subsequent analysis.
[0057] 2. Calculate text similarity and reconstruct the propagation network graph. A mature text segmentation system is used to segment the streaming text, and keywords and account information are extracted based on TF-IDF for subsequent analysis. Furthermore, to address issues such as semantic and contextual changes, advanced word vector technology is introduced to further improve the accuracy of the analysis.
[0058] Calculating text similarity: Using the word2vec algorithm, the semantic similarity between keywords in the text is calculated separately, and the edit distance of specific components (web page address, account) in the text is used to finally determine the text similarity between different texts. The calculation process is shown in the formula below.
[0059]
[0060] S_text represents the similarity between two texts, Text1 represents text 1, Text2 represents text 2, W_1,i represents words in text 1, W_2,j represents words in text 2, n1 represents the number of words in text 1, n2 represents the number of words in text 2, and Cos(w1i,w2j) represents the cosine distance between the two words. The word vector training set is a large-scale dataset containing positive and negative samples of fraudulent information, obtained through offline pre-training and updated regularly, significantly saving computation time.
[0061] Reconstruct the propagation network diagram, such as Figure 3 As shown:
[0062] The original propagation network graph is relatively sparse. A reconstructed propagation network graph is obtained by reconstructing it based on the number of similar text messages. This approach effectively counters anti-detection measures while allowing for manually controllable detection granularity, enabling dynamic adjustment of monitoring intensity based on different real-world situations. The original propagation network graph contains three independent propagation structures, where u1-u3 represent three sending accounts. Semantic analysis shows that the 13 text messages sent by u1-u3 have a text similarity exceeding a threshold; therefore, they are merged to obtain the reconstructed propagation network graph. During the merging process, the implicit relationships between different sending accounts are mined by calculating the text similarity of the text messages sent by the sending accounts. This is highly helpful in collecting and identifying positive and negative samples of fraudulent information.
[0063] 3. Fraud structural subgraph detection, PageRank rating, and fraud text information prototype extraction.
[0064] To further calculate the fraud weight value of the central node in the reconstructed propagation network graph, the PageRank algorithm is introduced to dynamically calculate the fraud weight value of the central node through the propagation relationship between different nodes.
[0065] To ensure that the PageRank algorithm's matrix calculation iterations converge.
[0066]
[0067] In the above formula, x is the transition matrix, α is the coefficient of the random jump factor, and |V| is the number of sending accounts in the central node of the reconstructed propagation network graph. The fraud weight value of the central node in the reconstructed propagation network graph comes partly from random walks generated based on real forwarding relationships, and partly from random jumps generated between similar text information. During information propagation, in addition to forwarding received messages, sending accounts may obtain messages through other means and then directly send information. The latter process corresponds to the random jump process. Using this method, the above process can guarantee convergence to a nontrivial solution.
[0068] Additionally, to implement this method, a Java runtime environment needs to be deployed and configured. Since the process is very simple and the same as installing ordinary computer software, it will not be described in detail here.
[0069] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
Claims
1. A method for detecting fraudulent information based on integrated propagation relationships, characterized in that, include: Obtain a first information group, a database of fraudulent accounts, and a database of legitimate accounts, wherein each piece of information in the first information group includes text information and the sending account; A second information group is determined from the first information group based on the fraud account database and the normal account database, wherein the sending account of each message in the second information group does not exist in either the fraud account database or the normal account database; Multiple target information groups are obtained based on the second information group, wherein the number of similar text information between the first sending account and the second sending account in each target information group is greater than a first threshold. Calculate the fraud weight value for each of the target information groups; If the fraud weight value of the target information group is greater than the second threshold, each text information in the target information group will be identified as fraud information. The method further includes, before obtaining multiple target information groups based on the second information group: Calculate the text similarity between the first text information and the second text information, wherein the first text information is the text information in the second information group, the second text information is the text information in the second information group excluding the first text information, and the first text information and the second text information are text information from different sending accounts; If the text similarity is greater than a third threshold, the first text information and the second text information are determined to be similar text information.
2. The method according to claim 1, characterized in that, The calculation of the text similarity between the first text information and the second text information includes: A first similarity value is determined based on the webpage addresses in the first and second text information; A second similarity value is determined based on the account information in the first and second text information; A third similarity value is determined based on the semantics of the first and second text information; The sum of the first similarity value, the second similarity value, and the third similarity value is determined as the text similarity.
3. The method according to claim 2, characterized in that, Determining the first similarity value based on the webpage addresses in the first text information and the second text information includes: Obtain the first webpage address from the first text information and the second webpage address from the second text information; If the first webpage address and the second webpage address are the same, the first similarity value is determined to be the first value; If the first webpage address and the second webpage address are different, the first similarity value is determined to be the second value.
4. The method according to claim 2, characterized in that, The step of determining the second similarity value based on the account in the first text information and the second text information includes: Obtain the first account from the first text information and the second account from the second text information; If the first account and the second account are the same, the second similarity value is determined to be the first value; If the first account and the second account are not the same, the second similarity value is determined to be the first value.
5. The method according to claim 2, characterized in that, The step of determining the third similarity value based on the semantics of the first text information and the second text information includes: The first text information is segmented into words to obtain a first word group, and the second text information is segmented into words to obtain a second word group; The third similarity value is calculated based on the first word group, the second word group, and the word embedding model.
6. The method according to claim 1, characterized in that, The calculation of the fraud weight value for each of the target information groups includes: Each of the target information groups is designated as the current target information group, and the following operations are performed on the current target information group: Determine the number of sending accounts in the current target information group; The fraud weight value of the current target information group is calculated based on the quantity and the random walk model.
Citation Information
Patent Citations
Risk group identification method and device, equipment and storage medium
CN115115370A