Fraud data analysis method and device, electronic equipment and storage medium
By classifying the main communication account name and wireless network name of the fraudulent data, and combining application installation information and chat data, the problem of low efficiency and low accuracy of fraudulent data analysis in the existing technology is solved, and fast and efficient fraudulent data analysis and high-accurate fraudulent information determination are achieved.
Patent Information
- Application Number
- CN202510061132.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-27
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is inefficient, cost-effective and low accuracy when analyzing fraudulent data, making it difficult to quickly and efficiently analyze fraudulent data in batches.
The fraud data is classified through the main communication account name and wireless network name, count the number of application installations to determine the fraudulent application, and determine the fraudulent account through the friend account and group member account, and finally determine the fraudulent information based on the chat data.
It realizes rapid and efficient batch analysis of fraudulent data, improves the accuracy of analysis results, and reduces labor costs.
Smart Images

Figure CN120067905A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data analysis, and particularly relates to a method, device, electronic device and storage medium for analyzing fraud data. Background Art
[0002] Telecom fraud refers to fabricating false information, setting up scams, and remotely and non-contactingly scamming defrauded accounts through telephone, network, and text message methods. In order to prevent more defrauded accounts from suffering losses due to telecom fraud, it is necessary to analyze fraud data to quickly discover valuable clues.
[0003] In the prior art, the method for analyzing fraud data is usually as follows: for each piece of fraud data, mining and analyzing it one by one from different dimensional contents to find fraud evidence or determine suspicious fraud accounts. However, this method of the prior art has the following problems: the time consumed for analyzing fraud data is relatively long, and the efficiency is low; a large amount of human cost needs to be invested, and the accuracy rate of the analysis result is relatively low. Summary of the Invention
[0004] The purpose of the present invention is to quickly and efficiently batch-analyze fraud data and improve the accuracy rate of the fraud data analysis result.
[0005] In a first aspect, an embodiment of the present invention provides a method for analyzing fraud data, the method comprising: Obtaining fraud data to be analyzed; For the first fraud data with a main communication account name in the fraud data to be analyzed, classifying the first fraud data based on the main communication account name, and determining the first fraud data associated with the main communication account name as the same type of first fraud data to obtain a first fraud data classification result, where the main communication account name is the account name corresponding to the account that actively initiates communication, and each type of first fraud data corresponds to a fraud group name; Obtaining the wireless network names corresponding to all the first fraud data in the first fraud data classification result, and updating the first fraud data classification result based on the obtained wireless network names and second fraud data to obtain a second fraud data classification result, where in the second fraud data classification result, the same type of fraud data corresponds to the same wireless network name, and the second fraud data is the fraud data other than the first fraud data in the fraud data to be analyzed; For each type of fraud data in the second fraud data classification result, obtaining the application program installation list corresponding to all the fraud data in each type of fraud data, counting the installation times of each application program in the application program installation list, and determining the social communication application program with the installation times greater than the preset times as the fraud application program; For each type of fraud data in the second fraud data classification result, determine the fraud-affected accounts based on the friend account names of each main communication account name and the group member account names of the group chats corresponding to each main communication account name. Based on the chat data of the fraud-affected accounts in the fraud application, determine the fraud information of the fraud-affected accounts.
[0006] Optionally, the classifying the first fraud data based on the main communication account name, determining the first fraud data belonging to the main communication account name as the same type of first fraud data, and obtaining the first fraud data classification result includes: Perform word segmentation on the main communication account name of each piece of first fraud data to obtain multiple word segmentation results. For each word segmentation result, retain the starting phrase in the word segmentation result to obtain multiple starting phrases. Count the number of times each starting phrase appears, and sort the starting phrases in descending order of the number of times the starting phrase appears to obtain a sorted list. In the order from the front to the back in the sorted list, sequentially determine the main communication account names that match each starting phrase. If multiple main communication account names match the same starting phrase, determine that the multiple main communication account names are associated, and determine the first fraud data corresponding to the multiple main communication account names as the same type of first fraud data to obtain the first fraud data classification result.
[0007] Optionally, the updating the first fraud data classification result based on the obtained wireless network names and the second fraud data to obtain the second fraud data classification result includes: Count the number of times each of the obtained wireless network names appears. Sort the wireless network names in descending order of the number of times the wireless network names appear to obtain a sorted list of wireless network names. In the order from the front to the back in the sorted list of wireless network names, sequentially determine whether the same type of first fraud data in the first fraud data classification result corresponds to the same wireless network name. If there is target fraud data in the same type of first fraud data that corresponds to different wireless network names, based on the first wireless network name corresponding to the target fraud data, adjust the target fraud data to the first fraud data category corresponding to the first wireless network name to obtain the updated first fraud data classification result as the first classification result. For the second fraud data, based on the second wireless network name corresponding to the second fraud data, the second fraud data is added to the first fraud data category corresponding to the second wireless network name, to obtain an updated first fraud data classification result as the second classification result; The first classification result and the second classification result are determined as a second fraud data classification result.
[0008] Optionally, for each type of fraud data in the second fraud data classification result, determining the defrauded account based on the friend account names of each main communication account name in each type of fraud data and the group member account names of the group chat corresponding to the main communication account name, includes: For each type of fraud data in the second fraud data classification result, traverse the friend account database based on the beginning phrase of each main communication account name in each type of fraud data, and determine the friend account name matching the beginning phrase as the target friend account name of the main communication account name, the friend account database is obtained based on the analysis of the friend account names corresponding to each main communication account name in the fraud data to be analyzed, and the main communication account name and the target friend account name correspond to the same fraud group name; The friend account names corresponding to the main communication account name, excluding the target friend account name, form a first account name set; The account names of the group members of the group chat corresponding to the main communication account name, except the main communication account and the target friend account name, are combined into a second account name set; An intersection of the first account name set and the second account name set is calculated to obtain identical account names included in the first account name set and the second account name set, and accounts corresponding to the identical account names are determined to be defrauded accounts.
[0009] Optionally, determining the fraudulent information of the victim account based on the chat data of the victim account in the fraudulent application includes: Obtaining private chat data and group chat data of the defrauded account in the fraud application; Extracting multiple keywords from the private chat data and the group chat data, and determining weights corresponding to the multiple keywords based on the number of times the multiple keywords appear respectively; The multiple keywords are sorted in descending order of weight to obtain a target number of target keywords with top rankings, and the fraudulent information of the defrauded account is determined based on the target keywords.
[0010] Optionally, determining the fraud-victimized information of the fraud-victimized account based on the target keyword includes: Match each communication record in the fraud data to be analyzed through the target keyword; If the target keyword matches the target communication record, determine the communication account corresponding to the target communication record as a defrauded account, and output a list of defrauded accounts; For each defrauded account in the list of defrauded accounts, extract the fraud information of the defrauded account based on the context communication content in the communication record of the defrauded account.
[0011] In a second aspect, an embodiment of the present invention provides a fraud data analysis device, the device includes: A fraud data acquisition module, configured to acquire fraud data to be analyzed; A first data classification module, for the first fraud data with a main communication account name in the fraud data to be analyzed, classify the first fraud data based on the main communication account name, and determine the first fraud data belonging to the main communication account name as the same type of first fraud data, to obtain a first fraud data classification result, the main communication account name is the account name corresponding to the account that actively initiates communication, and each type of first fraud data corresponds to a fraud group name; A second data classification module, configured to obtain the wireless network names corresponding to all the first fraud data in the first fraud data classification result, and update the first fraud data classification result based on the obtained wireless network names and the second fraud data, to obtain a second fraud data classification result, in the second fraud data classification result, the same type of fraud data corresponds to the same wireless network name, and the second fraud data is the fraud data other than the first fraud data in the fraud data to be analyzed; A fraud application determination module, for each type of fraud data in the second fraud data classification result, obtain the application installation list corresponding to all the fraud data in each type of fraud data, count the installation times of each application in the application installation list, and determine the social communication application with the installation times greater than the preset times as a fraud application; A defrauded account determination module, for each type of fraud data in the second fraud data classification result, determine the defrauded account based on the friend account names of each main communication account name and the group member account names of the group chats corresponding to each main communication account name; A defrauded information determination module, configured to determine the defrauded information of the defrauded account based on the chat data of the defrauded account in the fraud application.
[0012] In a third aspect, an embodiment of the present invention provides an electronic device, including: At least one processor; A memory for storing the at least one processor-executable instruction; Wherein, the at least one processor is configured to execute the instruction to implement the method described in the first aspect.
[0013] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute the method described in the first aspect.
[0014] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.
[0015] In the technical solution of the embodiment of the present invention, after obtaining the fraud data to be analyzed, first classify the first type of fraud data by the main communication account name to obtain the first fraud data classification result; then update the first fraud data classification result through the wireless network names of all the first fraud data in the first fraud data classification result, and add the second type of fraud data without a main communication account name to the corresponding first fraud data classification category based on the wireless network name to obtain the second fraud data classification result, wherein, in the second fraud data classification result, the same type of fraud data corresponds to the same wireless network name. Furthermore, the classification of all the fraud data to be analyzed is realized. Next, for each type of fraud data in the second fraud data classification result, the installation times of the application programs corresponding to the fraud data can be counted, and the social communication application program with more installation times is determined as the fraud application program, and this fraud application program is the fraud tool of the fraud user. After determining the fraud tool, then determine the defrauded accounts through the friend account names and group member account names of the main communication account name, and finally, determine the defrauded information of the defrauded accounts according to the chat data of the defrauded accounts in the fraud application program.
[0016] It can be seen that the technical solution of the embodiment of the present invention realizes the classification of all fraud data through the main communication account name and wireless network name of the fraud data to be analyzed, determines the fraud application program used by the fraud user through the application program installation information, and determines the defrauded accounts through the friend account names and group member account names, and finally accurately determines the defrauded information of the defrauded accounts according to the chat data of the defrauded accounts in the fraud application program. Thus, the batch analysis of fraud data is realized quickly and efficiently, and the accuracy rate of the analysis result is also relatively high. Description of the Drawings
[0017] Figure 1 It is a flowchart of a fraud data analysis method provided by an embodiment of the present invention; Figure 2 For Figure 1 the flowchart of the specific implementation of S120 in Figure 3 For Figure 1 the flowchart of the specific implementation of S130 in Figure 4 For Figure 1 the flowchart of the specific implementation of S150 in Figure 5 For Figure 1 the flowchart of the specific implementation of S160 in Figure 6 the structural schematic diagram of a fraud data analysis device provided by an embodiment of the present invention; Figure 7 the structural schematic diagram of an electronic device provided by an embodiment of the present invention. Specific implementation mode
[0018] The present invention will be described in detail below through embodiments.
[0019] Telecom fraud refers to fabricating false information, setting up scams through telephone, network, and SMS methods, and remotely and non-contactingly scamming the scammed accounts. In order to prevent more scammed accounts from suffering losses due to telecom fraud, it is necessary to analyze fraud data to quickly discover valuable clues.
[0020] In the prior art, the method for analyzing fraud data is usually: for each piece of fraud data, mining and analyzing it one by one from different dimensional contents to find fraud evidence or determine suspicious fraud accounts. However, this method of the prior art has the following problems: the time consumed for analyzing fraud data is relatively long, the efficiency is low; and a large amount of human cost needs to be invested, and the accuracy of the analysis result is relatively low.
[0021] In order to solve the above technical problems, the overall technical solution of the embodiment of the present invention is as follows: Step 1. After obtaining the exhibits, that is, the fraud data to be analyzed, perform the following data processing on the fraud data to establish a theme library.
[0022] (1) Establish a main communication account theme library. Aggregate the main communication account information and its relationships of all fraud data. Among them, the main account information can be the main communication account name, and the main communication account name is the account name corresponding to the account that actively initiates communication. And it is possible to analyze whether the starting phrases of multiple main communication account names are the same. If the starting phrases of multiple main communication account names are the same, it indicates that the company names packaged by multiple main communication account names are the same. Therefore, it can be determined that the accounts corresponding to multiple main communication account names are fraud accounts, and multiple fraud accounts belong to the same fraud gang, that is, multiple fraud accounts belong to the same fraud group.
[0023] (2) Establish a friend account library. Aggregate the friend account information and their relationships of all fraud data. The friend account information can be the friend account names, as well as the relationships between each friend account name. For example, whether the accounts corresponding to two friend account names are in the same group, etc.
[0024] (3) Establish a group member theme library. Aggregate the group member account information and their relationships of all fraud data. Among them, the group member account information can be the group member account names. For example, whether the accounts corresponding to two group member account names are in the same group, etc., and what is the account name of the main communication account of the group they are in, etc.
[0025] (4) Establish an APP installation theme library. Aggregate the APP installation list information and their relationships of all fraud data. For example, what are the APP names corresponding to all fraud data respectively, and which accounts have installed the same APP, etc.
[0026] (5) Establish a Wi-Fi connection theme library. Aggregate the Wi-Fi connection information and their relationships of all fraud data. For example, what are the Wi-Fi names corresponding to all fraud data respectively. Each fraud data corresponds to a Wi-Fi name, and this Wi-Fi name is the name of the Wi-Fi connected by the main communication account corresponding to this piece of fraud data. And the Wi-Fi connection theme library can also include: which fraud data have the same corresponding Wi-Fi name, etc.
[0027] Step 2: After obtaining the exhibits, i.e., the fraud data to be analyzed, the fraud data to be analyzed includes the first fraud data from mobile terminals such as mobile phones. Mobile phones usually install instant messaging tools, and fraud accounts use instant messaging tools to carry out fraud externally. The first fraud data includes the main communication account name, which is the account name corresponding to the account that actively initiates communication. Therefore, for the first fraud data in the fraud data to be analyzed, the main communication account name of the first fraud data can be analyzed, and the first fraud data is classified through the main communication account name. The first fraud data associated with the main communication account name is determined to be the same type of first fraud data, and the first fraud data classification result is obtained.
[0028] For example, when a fraudulent account commits fraud on a victimized account, it usually presents a formal company online. The fraudulent account serves as the main communication account, and the name of its main communication account is the first impression given to the victimized account when adding friends, suggesting that the victimized account is a legitimate company and gaining the trust of the victimized account. Therefore, the beginning of the main communication account name often contains the name of the company it purports to be, and the end is a personal nickname. Thus, by analyzing whether the beginnings of the main communication account names in the fraud data are similar company names, it is possible to determine whether the main communication account names are related. Specifically, if the beginnings of multiple main communication account names are similar company names, it can be preliminarily determined that the packaging companies corresponding to these multiple main communication account names are the same, that these main communication account names correspond to accounts belonging to the same fraud gang, and that the first fraud data corresponding to these multiple main communication account names is determined to be the same type of first fraud data.
[0029] Specifically, first perform word segmentation on all main communication account names to obtain multiple phrases, and retain the beginning phrases among the multiple phrases. Count the number of occurrences of the beginning phrases and sort them in descending order of the number of occurrences to obtain the result list K: K = {(K1, C1), (K2, C2), (K3, C3) …… (Kn, Cn)}, where (Ki, Ci) is the i-th result, Ki represents the beginning phrase, and Ci represents the number of occurrences of that beginning phrase.
[0030] First, use K1 to match the main communication account names in the main account theme library. If K1 can match multiple main communication account names, then it can be determined that the main communication accounts corresponding to these main communication account names represent the same fraud gang, that is, the same fraud group. The first fraud data corresponding to these main communication account names is classified as the same type of fraud data, and the group name of the fraud group to which these main communication account names and this type of fraud data belong is marked. This group name can also be called the gang name.
[0031] Second, use K2 to match the main communication account names in the main account theme library. When matching a main communication account name, determine whether it has already been marked with a group name. If it has already been marked, no further marking is performed; if it has not been marked with a group name, the main communication account name and the first fraud data corresponding to that main communication account name are separately marked with a new group name. Traverse K until all (Kn, Cn) are marked. Sort the marked group names to obtain T = {T1, T2, T3 …… Tn}, where Ti represents a single group name, which can be the group name corresponding to the main communication account names associated with the beginning phrase Ki.
[0032] Step 3. After classifying the first fraud data by the main communication account name to obtain the classification result of the first fraud data, the classification of the first fraud data with the main communication account name is realized, that is, the first fraud data have all been marked with the group name. However, for the second fraud data without the main communication account name in the data to be analyzed, the group name has not been marked yet. These second fraud data are mainly fraud data from desktop computers. Since desktop computers are generally used for internal management and office of fraud accounts and do not carry out fraud externally, and some instant messaging tools are not installed on desktop computers, it is impossible to classify the second fraud data by the main communication account name.
[0033] However, desktop computers generally also use wireless network cards to connect to Wifi for Internet access and connect to the same Wifi as mobile phones. At this time, classification can be carried out using the Wifi name. The second fraud data without the main account communication name in the fraud data to be analyzed can also be classified according to the wireless network names corresponding to all the first fraud data in the classification result of the first fraud data, so as to realize the classification of all the fraud data to be analyzed and obtain the fraud gangs corresponding to each type of fraud data, that is, obtain the fraud groups corresponding to each type of fraud data.
[0034] Specifically, the classification of all fraud data can be realized through the following three steps.
[0035] The first step is to obtain the list of Wifi names of all the first fraud data in the classification result of the first fraud data, perform aggregation statistics according to the Wifi names, and obtain the number of times each Wifi name appears. If the number of times a Wifi name appears is 1, it means that only one mobile terminal is connected to the Wifi with this Wifi name, and the probability that this Wifi name is connected by a fraud gang is relatively low. Therefore, the Wifi names that appear only once are excluded. The remaining Wifi names are sorted in descending order of the number of appearances to obtain the Wifi name list W = {W1, W2, W3... Wn}, where Wi represents the i-th Wifi.
[0036] The second step is, after obtaining the Wifi name list, for any Wifi name in the Wifi name list, the Wifi name can be associated with the fraud data corresponding to the Wifi name. Specifically, if the wireless network names corresponding to the first fraud data belonging to the group name Ti are the same, then there is no need to adjust the first fraud data belonging to the group name Ti; if the wireless network names corresponding to the first fraud data belonging to the group name Ti have different wireless network names, then the first fraud data corresponding to the different wireless network names need to be deleted from the first fraud data of the group name Ti, and the first fraud data is adjusted to the first fraud data category corresponding to its wireless network name, for example, adjusted to the first fraud data category with the group name Tj, indicating that the same fraud group may have internal division of labor to package multiple company names for fraud, thereby achieving the update of the first fraud data classification result, and determining the updated first fraud data classification result as the second fraud data classification result.
[0037] For the second fraud data in which the main communication account name does not exist in the fraud data to be analyzed, the Wifi name corresponding to each second fraud data can be obtained in the Wifi connection theme library, and the second fraud data can be adjusted to the first fraud data category corresponding to its Wifi name, thereby updating the first fraud data classification result and determining the updated first fraud data classification result as the second fraud data classification result.
[0038] In the third step, repeat the first and second steps until the first fraud data corresponding to any Ti in T is traversed. At this time, all the fraud data to be analyzed have been labeled with the group name, that is, all the fraud data to be analyzed have been classified to obtain different types of fraud data sets, and each type of fraud data set corresponds to a group name.
[0039] From the third step above, we can know that each type of fraud data set corresponds to a group name, and each type of fraud data set corresponds to one or more main communication account name initial phrases. Therefore, the correspondence between the main communication account name initial phrases and the group name can be determined. In addition, the correspondence between the main communication account name initial phrases and the group name can be expressed as: R={(K1,T1),(K2,T2),(K3,T3)……(Kn,Tn)}, where (Ki,Ti) represents the i-th correspondence, Ki is the i-th main communication account name initial phrase, and Ti is the corresponding group name. In actual applications, there may be a situation where multiple Ks correspond to the same T, that is, the situation where multiple main communication account name initial phrases correspond to one group name.
[0040] Step 4: Uncover the methods and tools used by fraudulent accounts.
[0041] Traverse all fraud data corresponding to group name T. Specifically, obtain the APP installation list of all fraud data corresponding to the i-th group name Ti in the APP installation theme library. In order to effectively analyze which APPs the fraud account uses to commit fraud on the victim account, after obtaining the APP installation list, the APPs that come with the terminal system can be eliminated. Among the remaining APPs, the number of installations of each APP is aggregated and counted according to the software name, and the APPs are sorted from large to small according to the number of APP installations to obtain a sorted list A={A1,A2,A3,Ai,...An}, where Ai represents any APP, and is matched with the APP knowledge base configured in the system to obtain the category corresponding to each APP. The top-ranked social instant messaging APPs are basically fraud tools used by fraud users to communicate with the victim account. Tool APPs such as translation software, VPN (Virtual Private Network), customer service APPs, etc. are basically auxiliary tools. There are also some software similar to ERP (Enterprise Resource Planning) that may be used for internal management and can be further analyzed and verified.
[0042] Step 5: Withdraw the defrauded account.
[0043] (1) Gather friend account names. According to the correspondence R={(K1,T1),(K2,T2),(K3,T3)…(Kn,Tn)} between the initial phrases of the main communication account names and the group names calculated in step 3, traverse the friend account database through the initial phrase Ki of the main communication account names. If a friend account name matches Ki, such as the initial phrase of the friend account name is the same as Ki, then the group name is marked as Ti for the friend account name. In other words, the friend account corresponding to the friend account name belongs to the same fraud group as the main communication account, that is, both of them enter the same fraud gang. The main communication account here refers to the account corresponding to the main communication account name with the initial phrase Ki. Through this step, the friend account belonging to the same gang as the main communication account can be found. As gang accounts, both the main communication account and the gang account are fraud accounts.
[0044] (2) According to the correspondence relationship R={(K1,T1),(K2,T2),(K3,T3)……(Kn,Tn)} between the starting phrase of the main communication account name and the group name calculated in step 3, traverse the group member theme library through the starting phrase Ki of the main communication account name to obtain the group member accounts of the group where the main account name with the starting phrase Ki is located. Among them, the group member accounts include the gang accounts of the main communication account and also include the defrauded accounts. Because when the fraud account implements fraud on the defrauded account, the fraud account usually first adds the defrauded account as a friend and then adds the defrauded account to the gang group to carry out telecommunications fraud on the fraud account in the gang group. Therefore, in addition to being in the group member accounts, the defrauded account is also the friend account of the fraud account.
[0045] (3) After obtaining the gang accounts of the same gang as the main communication account in (2), the set of friend account names F={F1,F2,F3……Fn} after removing the gang accounts can be obtained, where Fi represents the i-th friend account, and the set of group member account names M={M1,M2,M3……Mn} after removing the gang accounts, where Mi is the i-th group member. Next, find the intersection of F and M, that is, F∩M, to obtain the accounts that are both the friend accounts of the fraud accounts and the group member accounts. These accounts are the defrauded accounts.
[0046] (4) Extract the private chat data and group chat data of the fraud application of the above-mentioned defrauded accounts in step 4, and perform word segmentation on the private chat data and group chat data respectively to obtain multiple segmented word groups. Count the occurrence times of the multiple segmented word groups respectively, and calculate the word frequencies corresponding to each segmented word group. Remove the common words, and the keywords with higher occurrence frequencies are obtained. The keywords with higher occurrence frequencies are the keywords with higher weights. Or, the keywords with higher weights can be statistically obtained from the multiple segmented word groups through the TF-IDF algorithm. Of course, in actual applications, other methods can also be used to obtain the weights of each keyword, which will not be listed one by one here. Assume that the keyword set of the defrauded account is V={V1,V2,V3……Vn}, where Vi represents the i-th keyword.
[0047] Since in the actual fraud application scenario, the communication process between the fraud account and the defrauded account is basically the same as the script, the keywords extracted from the chat data are basically the same as the defrauded situation. By performing context association on the keywords with high weights in the chat data, the key points of the defrauded account being defrauded can be restored.
[0048] (5) Use the keywords V = {V1, V2, V3... Vn} of the chat data of the defrauded accounts that have been extracted to batch match the communication records in the fraud data to be analyzed. All the communication accounts corresponding to the communication records that hit the keywords may be defrauded accounts, so that the defrauded accounts can be identified, and a list of defrauded account numbers can be output. Then, contact the context of the communication records and extract information such as the defrauded amount from the communication records. Among them, the communication records refer to the behavior records and content records of communicating with each other through the Internet and the telecommunications network, including various forms such as making phone calls, sending text messages, chat data, and e-mails.
[0049] In the technical solution of the embodiment of the present invention, after obtaining the fraud data to be analyzed, first classify the first type of fraud data by the main communication account name to obtain the first fraud data classification result; then update the first fraud data classification result by the Wifi names of all the first fraud data in the first fraud data classification result, and add the second type of fraud data without the main communication account name to the corresponding first fraud data classification category based on the Wifi name to obtain the second fraud data classification result. Among them, in the second fraud data classification result, the same type of fraud data corresponds to the same wireless network name. Thus, the classification of all the fraud data to be analyzed is realized. Next, for each type of fraud data in the second fraud data classification result, the installation times of the application programs corresponding to the fraud data can be counted, and the social communication application program with more installation times can be determined as the fraud application program, and this fraud application program is the fraud tool of the fraud user. After determining the fraud tool, then determine the defrauded accounts through the friend account names and group member account names of the main communication account name. Finally, according to the chat data of the defrauded accounts in the fraud application program, determine the defrauded information of the defrauded accounts.
[0050] It can be seen that in the technical solution of the embodiment of the present invention, all the fraud data is classified by the main communication account name and the Wifi name of the fraud data to be analyzed, the fraud application program used by the fraud user is determined through the application program installation information, the defrauded accounts are determined through the friend account names and group member account names, and finally the defrauded information of the defrauded accounts is accurately determined through the chat data of the defrauded accounts in the fraud application program. Thus, the batch analysis of fraud data is realized quickly and efficiently, and the accuracy rate of the analysis result is also relatively high.
[0051] After elaborating on the overall technical solution of this solution in detail, the following will elaborate on a fraud data analysis method provided by the embodiment of the present invention.
[0052] As Figure 1 shown, the embodiment of the present invention provides a fraud data analysis method, and this method may include the following steps: S110, Obtain fraud data to be analyzed.
[0053] Among them, the fraud data to be analyzed can be roughly divided into two categories. One is the first type of fraud data with a main communication account name, and the other is the second type of fraud data without a main communication account name.
[0054] S120, For the first type of fraud data with a main communication account name in the fraud data to be analyzed, classify the first type of fraud data based on the main communication account name, and determine the first type of fraud data associated with the main communication account name as the same type of first type of fraud data, to obtain the classification result of the first type of fraud data.
[0055] Among them, the main communication account name is the account name corresponding to the account that actively initiates communication, and each type of first type of fraud data corresponds to a fraud group name.
[0056] Specifically, when a fraud account implements a fraud behavior on a defrauded account, it generally packages a formal company on the network. The fraud account is the main communication account, and the main communication account name of the fraud account is the first impression given to the defrauded account when adding a friend, suggesting that the defrauded account is a regular company and obtaining the trust of the defrauded account. Therefore, the beginning of the main communication account name is often the company name of its packaged company, and the end of the main communication account name is the personal nickname. Then, by analyzing whether the beginnings of the main communication account names of the fraud data are similar company names, it can be judged whether the main communication account names are associated. If multiple main communication account names are associated, it means that the accounts corresponding to these multiple main communication account names are the same fraud gang, and the first type of fraud data associated with the main communication account name can be determined as the same type of first type of fraud data.
[0057] Specifically, in one implementation, as Figure 2 shown, S120, classify the first type of fraud data based on the main communication account name, determine the first type of fraud data associated with the main communication account name as the same type of first type of fraud data, to obtain the classification result of the first type of fraud data, which can include the following steps: S121, Perform word segmentation on the main communication account name of each piece of the first type of fraud data to obtain multiple word segmentation results.
[0058] S122, For each word segmentation result, retain the beginning phrase in the word segmentation result to obtain multiple beginning phrases.
[0059] S123, Count the number of occurrences of each beginning phrase, and sort the beginning phrases in descending order according to the number of occurrences of the beginning phrases to obtain a sorted list.
[0060] S124. Determine the main communication account names that match each starting phrase in sequence according to the order from the front to the back in the sorted list.
[0061] S125. If multiple main communication account names match the same starting phrase, determine that the multiple main communication account names are associated, and determine the first fraud data corresponding to the multiple main communication account names as the same type of first fraud data to obtain the first fraud data classification result.
[0062] It should be noted that the specific implementation manner of S120 has been elaborated in detail in the overall technical solution of the present invention and will not be repeated here.
[0063] S130. Obtain the wireless network names corresponding to all the first fraud data in the first fraud data classification result, and update the first fraud data classification result based on the obtained wireless network names and the second fraud data to obtain the second fraud data classification result.
[0064] Among them, in the second fraud data classification result, the same type of fraud data corresponds to the same wireless network name, and the second fraud data is the fraud data other than the first fraud data in the fraud data to be analyzed.
[0065] Specifically, after classifying the first fraud data by the main communication account name to obtain the first fraud data classification result, the classification of the first fraud data with the main communication account name is realized, that is, the first fraud data has been marked with the group name. However, for the second fraud data without the main communication account name in the data to be analyzed, the group name has not been marked. These second fraud data are mainly fraud data from desktop computers. Since desktop computers are generally used for internal management and office of fraud accounts and do not carry out fraud externally, and some instant messaging tools are not installed on desktop computers, the second fraud data cannot be classified by the main communication account name.
[0066] However, desktop computers generally also use wireless network cards to connect to Wifi to access the Internet, and connect to the same Wifi as mobile phones. At this time, classification can be carried out using the Wifi name. The second fraud data without the main account communication name in the fraud data to be analyzed can also be classified according to the wireless network names corresponding to all the first fraud data in the first fraud data classification result, so as to realize the classification of all the fraud data to be analyzed, obtain the fraud gang corresponding to each type of fraud data, that is, obtain the fraud group corresponding to each type of fraud data. In addition, the same fraud group connects to the same Wifi. Therefore, the first fraud data classification result can be updated using the Wifi name to obtain the second fraud data classification result, and the same type of fraud data in the second fraud data classification result has the same Wifi name.
[0067] In one embodiment, S130, update the first fraud data classification result based on the obtained wireless network names and the second fraud data to obtain the second fraud data classification result. As Figure 3 shown, the following steps may be included: S131, count the number of occurrences of each of the obtained wireless network names.
[0068] S132, sort the wireless network names in descending order of the number of occurrences of the wireless network names to obtain a sorted list of wireless network names; S133, sequentially determine whether the same type of first fraud data in the first fraud data classification result corresponds to the same wireless network name in the order from the front to the back of the sorted list of wireless network names; S134, if there is target fraud data in the same type of first fraud data that corresponds to different wireless network names, based on the first wireless network name corresponding to the target fraud data, adjust the target fraud data to the first fraud data category corresponding to the first wireless network name to obtain the updated first fraud data classification result as the first classification result.
[0069] S135, for the second fraud data, based on the second wireless network name corresponding to the second fraud data, add the second fraud data to the first fraud data category corresponding to the second wireless network name to obtain the updated first fraud data classification result as the second classification result.
[0070] S136, determine the first classification result and the second classification result as the second fraud data classification result.
[0071] It should be noted that the specific implementation manner of S130 has been elaborated in detail in the overall technical solution of the present invention and will not be repeated here.
[0072] S140, for each type of fraud data in the second fraud data classification result, obtain the application installation list corresponding to all fraud data in each type of fraud data, count the installation times of each application in the application installation list, and determine the social communication applications with installation times greater than the preset times as fraud applications.
[0073] Specifically, an APP installation list of all fraudulent data corresponding to the i-th group name Ti is obtained in the APP installation theme library. In order to effectively analyze which APPs are used by the fraudulent account to commit fraud on the defrauded account, after obtaining the APP installation list, the APPs that come with the terminal system can be eliminated. Among the remaining APPs, the number of installations of each APP is aggregated and counted according to the software name, and the APPs are sorted from large to small according to the number of APP installations to obtain a sorted list A={A1, A2, A3, Ai, ...An}, where Ai represents any APP, and is matched with the APP knowledge base configured in the system to obtain the category corresponding to each APP. The top-ranked social instant messaging APPs are basically fraud tools used by fraudulent users to communicate with the defrauded account. Therefore, social communication applications with more than a preset number of installations are determined as fraudulent applications. The preset number can be limited according to actual conditions, and the embodiments of the present invention do not make specific limitations on this.
[0074] S150, for each type of fraud data in the second fraud data classification result, determine the defrauded account based on the friend account names of each main communication account name in each type of fraud data and the group member account names of the group chat corresponding to each main communication account name.
[0075] In one embodiment, S150, for each type of fraud data in the second fraud data classification result, based on the friend account names of each main communication account name in each type of fraud data and the group member account names of the group chat corresponding to the main communication account name, determine the fraudulent account, such as Figure 4 As shown, including: S151, for each type of fraud data in the second fraud data classification result, traverse the friend account database based on the beginning phrase of each main communication account name in each type of fraud data, and determine the friend account name that matches the beginning phrase as the target friend account name of the main communication account name.
[0076] Among them, the friend account database is obtained based on the analysis of the friend account names of each main communication account name in the analyzed fraud data, and the main communication account name and the target friend account name correspond to the same fraud group name.
[0077] S152: assemble the friend account names corresponding to the main communication account name, excluding the target friend account name, into a first account name set.
[0078] S153: Group the account names of the group members of the group chat corresponding to the main communication account name, except the main communication account name and the target friend account name, into a second account name set.
[0079] S154. Calculate the intersection of the first account name set and the second account name set to obtain the same account names of the first account set and the second account set, and determine the accounts corresponding to the same account names as the defrauded accounts. Specifically, for each type of fraud data in the second fraud data classification result, traverse the friend account library through the starting phrase Ki of the main communication account name. If a friend account name matches Ki, that is, the starting phrase of the friend account name is the same as Ki, then label the group name of the friend account name as Ti. That is to say, the friend account corresponding to the friend account name belongs to the same fraud group as the main communication account, that is, both are input into the same fraud gang. Here, the main communication account refers to the account corresponding to the main communication account name with the starting phrase Ki. Through this step, the friend accounts belonging to the same gang as the main communication account can be found and used as gang accounts. Both the main communication account and the gang accounts belong to the defrauded accounts.
[0080] Then, traverse the group member theme library through the starting phrase Ki of the main communication account name to obtain the group member accounts of the group where the main account name with the starting phrase Ki is located. Among them, the group member accounts include the gang accounts of the main communication account and also the defrauded accounts. Because when the fraud account implements fraud on the defrauded account, the fraud account usually first adds the defrauded account as a friend and then adds the defrauded account to the gang group to conduct telecommunications fraud on the defrauded account in the gang group. Therefore, the defrauded account is not only in the group member accounts but also a friend account of the fraud account.
[0081] After obtaining the gang account names that are in the same gang as the main communication account, the set of friend account names after removing the gang account names can be obtained as the first account name set F, F = {F1, F2, F3... Fn}, where Fi represents the account name of the i-th friend account, and the set of group member account names after removing the gang accounts can be obtained as the second account name set M, M = {M1, M2, M3... Mn}, where Mi is the account name of the i-th group member account. Next, find the intersection of F and M, that is, F ∩ M, to obtain the accounts that are both friend accounts of the fraud account and group member accounts. These accounts are the defrauded accounts.
[0082] S160. Based on the chat data of the defrauded account in the fraud application, determine the fraud information of the defrauded account.
[0083] Specifically, after determining the fraud application and the defrauded account, the chat data of the defrauded account in the fraud application can be obtained. These chat data include private chat data and group chat data with the fraud account. Based on these private chat data and group chat data, the fraud information of the defrauded account can be analyzed. Among them, the fraud information can include the amount of fraud, etc.
[0084] For the sake of clear description of the solution, the specific implementation manner of S160 will be elaborated in detail in the following embodiments.
[0085] In the technical solution of the embodiment of the present invention, after obtaining the fraud data to be analyzed, first, classify the first type of fraud data by the main communication account name to obtain the first fraud data classification result; then, update the first fraud data classification result by the wireless network names of all the first fraud data in the first fraud data classification result, and add the second type of fraud data without the main communication account name to the corresponding first fraud data classification category based on the wireless network name to obtain the second fraud data classification result, where in the second fraud data classification result, the same type of fraud data corresponds to the same wireless network name. Thus, the classification of all the fraud data to be analyzed is realized. Next, for each type of fraud data in the second fraud data classification result, the installation times of the application programs corresponding to the fraud data can be counted, and the social communication application program with more installation times is determined as the fraud application program, and this fraud application program is the fraud tool of the fraud user. After determining the fraud tool, then determine the fraud-affected accounts through the friend account names and group member account names of the main communication account name, and finally, determine the fraud information of the fraud-affected accounts according to the chat data of the fraud-affected accounts in the fraud application program.
[0086] It can be seen that in the technical solution of the embodiment of the present invention, the classification of all fraud data is realized by the main communication account name and wireless network name of the fraud data to be analyzed, the fraud application program used by the fraud user is determined through the application program installation information, and the fraud-affected accounts are determined through the friend account names and group member account names. Finally, the fraud information of the fraud-affected accounts is accurately determined according to the chat data of the fraud-affected accounts in the fraud application program. Thus, the rapid and efficient batch analysis of fraud data is realized, and the accuracy rate of the analysis result is also relatively high.
[0087] For the sake of clear description of the solution, the specific implementation manner of S160 will be elaborated in detail in the following embodiments.
[0088] In one implementation manner, S160 determines the fraud information of the fraud-affected accounts based on the chat data of the fraud-affected accounts in the fraud application program. As Figure 5 shown, it may include the following steps: S161, obtain the private chat data and group chat data of the fraud-affected accounts in the fraud application program.
[0089] S162, extract multiple keywords from the private chat data and group chat data, and determine the weights corresponding to the multiple keywords respectively based on the number of times each keyword appears.
[0090] S163. Sort multiple keywords in descending order of weight to obtain the target number of target keywords with higher rankings, and determine the fraud information of the fraud-affected account based on the target keywords.
[0091] Specifically, after determining the fraud-affected account and the fraud application, the private chat data and group chat data of the fraud-affected account in the fraud application can be extracted from the fraud data to be analyzed, and the private chat data and group chat data are respectively segmented to obtain multiple segmented word groups. The occurrence times of the multiple segmented word groups are respectively counted, and the word frequencies corresponding to each segmented word group are calculated. After removing common words, the keywords with higher occurrence frequencies are obtained. The keywords with higher occurrence frequencies are the keywords with higher weights. Alternatively, the keywords with relatively higher weights can be statistically obtained from multiple segmented word groups through the TF-IDF algorithm. Of course, in actual applications, other methods can also be used to obtain higher weights for each keyword, which will not be listed one by one here. Assume that the keyword set of the fraud-affected account is V = {V1, V2, V3... Vn}, where Vi represents the i-th keyword. The first N keywords in the keyword set can be determined as the target keywords, and N is the target number, which can be set according to the actual situation. The embodiments of the present invention do not make specific limitations on this.
[0092] Since in the actual fraud application scenario, the communication process between the fraud account and the fraud-affected account is basically the same as the script, the keywords extracted from the chat data are basically the same as the fraud situation. By performing context association on the keywords with high weights in the chat data, the fraud information of the fraud-affected account can be restored. The fraud information can include the fraud time and the fraud amount, etc.
[0093] It can be seen that in this embodiment, by determining the fraud-affected account and the fraud application, and extracting the chat data of the fraud-affected account in the fraud application, the fraud information of the fraud-affected account can be quickly and accurately obtained.
[0094] As an implementation manner of the embodiments of the present disclosure, determining the fraud information of the fraud-affected account based on the target keywords may include the following steps, namely steps a1 to a3: Step a1. Match each communication record in the fraud data to be analyzed through the target keywords.
[0095] Step a2. If the target keyword matches the target communication record, determine the communication account corresponding to the target communication record as the fraud-affected account, and output a list of fraud-affected accounts.
[0096] Step a3. For each fraud-affected account in the list of fraud-affected accounts, extract the fraud information of the fraud-affected account based on the context communication content in the communication record of the fraud-affected account.
[0097] Specifically, after extracting the keywords V = {V1, V2, V3... Vn} of the chat data of the defrauded account, the communication records in the fraud data to be analyzed can be batch-matched. All communication accounts corresponding to the communication records that hit the keywords may be defrauded accounts, so that the defrauded accounts can be identified and the list of defrauded account numbers can be output. Then, analyze the context of the communication records that hit the keywords, and extract information such as the defrauded amount from the communication records.
[0098] It can be seen that the technical solution of this embodiment can determine more defrauded accounts by analyzing the keywords of the chat data of the defrauded account, and quickly and accurately determine the defrauded information of each defrauded account.
[0099] In a second aspect, an embodiment of the present invention provides a fraud data analysis device 60, as Figure 6 shown. The device includes: A fraud data acquisition module 610, configured to acquire fraud data to be analyzed; A first data classification module 620, configured to classify the first fraud data with a main communication account name in the fraud data to be analyzed based on the main communication account name, and determine the first fraud data belonging to the main communication account name as the same type of first fraud data to obtain a first fraud data classification result. The main communication account name is the account name corresponding to the account that actively initiates communication, and each type of first fraud data corresponds to a fraud group name; A second data classification module 630, configured to acquire the wireless network names corresponding to all the first fraud data in the first fraud data classification result, and update the first fraud data classification result based on the acquired wireless network names and the second fraud data to obtain a second fraud data classification result. In the second fraud data classification result, the same type of fraud data corresponds to the same wireless network name. The second fraud data is the fraud data other than the first fraud data in the fraud data to be analyzed; A fraud application program determination module 640, configured to, for each type of fraud data in the second fraud data classification result, acquire the application program installation list corresponding to all the fraud data in each type of fraud data, count the installation times of each application program in the application program installation list, and determine the social communication application program with the installation times greater than the preset times as the fraud application program; A defrauded account determination module 650, configured to, for each type of fraud data in the second fraud data classification result, determine the defrauded account based on the friend account names of each main communication account name and the group member account names of the group chats corresponding to each main communication account name; A fraud information determination module 660, configured to determine fraud information of the fraud-affected account based on chat data of the fraud-affected account in the fraud application.
[0100] In the technical solution of the embodiment of the present invention, after obtaining the fraud data to be analyzed, first, the main communication account name is used to classify the first type of fraud data to obtain a first fraud data classification result; then, the wireless network name of all the first fraud data in the first fraud data classification result is used to update the first fraud data classification result, and the second type of fraud data without the main communication account name is also added to the corresponding first fraud data classification category based on the wireless network name to obtain a second fraud data classification result, where in the second fraud data classification result, the same type of fraud data corresponds to the same wireless network name. Furthermore, classification of all the fraud data to be analyzed is realized. Next, for each type of fraud data in the second fraud data classification result, the installation times of the application corresponding to the fraud data can be counted, and the social communication application with a larger installation times is determined as the fraud application, and this fraud application is the fraud tool of the fraud user. After determining the fraud tool, the fraud-affected account is determined by the friend account name of the main communication account name and the group member account names in the group where the fraud-affected account is located. Finally, the fraud information of the fraud-affected account is determined according to the chat data of the fraud-affected account in the fraud application.
[0101] It can be seen that in the technical solution of the embodiment of the present invention, classification of all the fraud data is realized by the main communication account name and the wireless network name of the fraud data to be analyzed, the fraud application used by the fraud user is determined through the application installation information, the fraud-affected account is determined by the friend account name and the group member account name, and finally, the fraud information of the fraud-affected account is accurately determined according to the chat data of the fraud-affected account in the fraud application. Thus, rapid and efficient batch analysis of fraud data is realized, and the accuracy rate of the analysis result is relatively high.
[0102] Optionally, the first data classification module is specifically configured to: Perform word segmentation on the main communication account name of each piece of the first fraud data to obtain a plurality of word segmentation results; For each word segmentation result, retain the starting phrase in the word segmentation result to obtain a plurality of starting phrases; Count the number of times each starting phrase appears, and sort the starting phrases in descending order of the number of times the starting phrase appears to obtain a sorted list; Determine the main communication account name that matches each starting phrase in turn according to the order from the front to the back in the sorted list; If multiple primary communication account names match the same starting phrase, determine that the multiple primary communication account names are associated, and determine the first fraud data corresponding to the multiple primary communication account names as the same type of first fraud data, obtaining a first fraud data classification result.
[0103] Optionally, the second data classification module is specifically configured to: Count the number of times each of the obtained all wireless network names appears; Sort the wireless network names in descending order of the number of times they appear, obtaining a sorted list of wireless network names; Sequentially determine whether the same type of first fraud data in the first fraud data classification result corresponds to the same wireless network name in the order from the front to the back of the sorted list of wireless network names; If there is target fraud data in the same type of first fraud data that corresponds to different wireless network names, based on the first wireless network name corresponding to the target fraud data, adjust the target fraud data to the first fraud data category corresponding to the first wireless network name, obtaining an updated first fraud data classification result as the first classification result; For the second fraud data, based on the second wireless network name corresponding to the second fraud data, add the second fraud data to the first fraud data category corresponding to the second wireless network name, obtaining an updated first fraud data classification result as the second classification result; Determine the first classification result and the second classification result as the second fraud data classification result.
[0104] Optionally, the defrauded account determination module is specifically configured to: For each type of fraud data in the second fraud data classification result, traverse the friend account database based on the starting phrase of each primary communication account name in each type of fraud data, and determine the friend account name that matches the starting phrase as the target friend account name of the primary communication account name. The friend account database is obtained by analyzing the friend account names corresponding to each primary communication account name in the fraud data to be analyzed, and the fraud group names corresponding to the primary communication account name and the target friend account name are the same; Form a first account name set from the friend account names corresponding to the primary communication account name except the target friend account name; Form a second account name set from the account names of the group members in the group chat corresponding to the primary communication account name except the primary communication account name and the target friend account name; An intersection of the first account name set and the second account name set is calculated to obtain identical account names included in the first account name set and the second account name set, and accounts corresponding to the identical account names are determined to be defrauded accounts.
[0105] Optionally, the fraudulent information determination module is specifically used to: Obtaining private chat data and group chat data of the defrauded account in the fraud application; Extracting multiple keywords from the private chat data and the group chat data, and determining weights corresponding to the multiple keywords based on the number of times the multiple keywords appear respectively; The multiple keywords are sorted in descending order of weight to obtain a target number of target keywords with top rankings, and the fraudulent information of the defrauded account is determined based on the target keywords.
[0106] Optionally, the fraudulent information determination module is further specifically used to: Matching each communication record in the fraud data to be analyzed by using the target keyword; If the target keyword matches the target communication record, determining that the communication account corresponding to the target communication record is a fraudulent account, and outputting a list of fraudulent accounts; For each defrauded account in the defrauded account list, defrauded information of the defrauded account is extracted based on the contextual communication content in the communication record of the defrauded account.
[0107] In a third aspect, an embodiment of the present invention provides an electronic device 700, such as Figure 7 As shown, including: at least one processor 701; a memory 702 for storing the at least one processor executable instruction; The at least one processor is configured to execute the instructions to implement the method described in the first aspect.
[0108] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method described in the first aspect.
[0109] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0110] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention without departing from the principles and spirit of the present invention.
Claims
1. A fraud data analysis method, characterized in that: The method comprises: Obtain fraud data to be analyzed; For the first fraud data in which the main communication account name exists in the fraud data to be analyzed, the first fraud data is classified based on the main communication account name, and the first fraud data associated with the main communication account name is determined to be the same type of first fraud data, to obtain the first fraud data classification result, wherein the main communication account name is the account name corresponding to the account that actively initiates the communication, and each type of first fraud data corresponds to a fraud group name; Obtaining wireless network names corresponding to all first fraudulent data in the first fraudulent data classification result, and updating the first fraudulent data classification result based on the obtained wireless network name and the second fraudulent data to obtain a second fraudulent data classification result, wherein in the second fraudulent data classification result, fraudulent data of the same type corresponds to the same wireless network name, and the second fraudulent data is the fraudulent data to be analyzed except the first fraudulent data; For each type of fraud data in the second fraud data classification result, obtain an application installation list corresponding to all fraud data in each type of fraud data, count the number of installations of each application in the application installation list, and determine a social communication application with a number of installations greater than a preset number as a fraud application; For each type of fraud data in the second fraud data classification result, based on the friend account names of each main communication account name in each type of fraud data and the group member account names of the group chat corresponding to each main communication account name, determine the defrauded account; Defrauded information of the defrauded account is determined based on the chat data of the defrauded account in the fraud application.
2. The method according to claim 1, characterized in that The first fraud data is classified based on the name of the main communication account, and the first fraud data associated with the name of the main communication account is determined to be the same type of first fraud data, and the first fraud data classification result is obtained, including: Segment the name of the main communication account of each first fraud data to obtain multiple segmentation results; For each word segmentation result, retain the beginning phrase in the word segmentation result to obtain multiple beginning phrases; Counting the number of occurrences of each opening phrase, and sorting the opening phrases in descending order of the number of occurrences of the opening phrases to obtain a sorted list; According to the order from the front to the back in the sorted list, the main communication account name matching each initial phrase is determined in turn; If multiple main communication account names match the same beginning phrase, it is determined that the multiple main communication account names are associated, and the first fraud data corresponding to the multiple main communication account names are determined as the same type of first fraud data to obtain a first fraud data classification result.
3. The method according to claim 1, characterized in that: The updating of the first fraud data classification result based on the acquired wireless network name and the second fraud data to obtain the second fraud data classification result includes: Count the number of occurrences of all acquired wireless network names; Sort the wireless network names in descending order of the number of times the wireless network names appear, to obtain a sorted list of wireless network names; According to the order of the wireless network name sorting list from front to back, it is determined in sequence whether the same type of first fraudulent data in the first fraudulent data classification result corresponds to the same wireless network name; If there is target fraud data corresponding to a different wireless network name in the same type of first fraud data, based on the first wireless network name corresponding to the target fraud data, the target fraud data is adjusted to the first fraud data category corresponding to the first wireless network name, and an updated first fraud data classification result is obtained as the first classification result; For the second fraud data, based on the second wireless network name corresponding to the second fraud data, the second fraud data is added to the first fraud data category corresponding to the second wireless network name, to obtain an updated first fraud data classification result as the second classification result; The first classification result and the second classification result are determined as a second fraud data classification result.
4. The method according to claim 1, characterized in that: For each type of fraud data in the second fraud data classification result, based on the friend account names of each main communication account name in each type of fraud data and the group member account names of the group chat corresponding to the main communication account name, determining the fraudulent account includes: For each type of fraud data in the second fraud data classification result, traverse the friend account database based on the beginning phrase of each main communication account name in each type of fraud data, and determine the friend account name matching the beginning phrase as the target friend account name of the main communication account name, the friend account database is obtained based on the analysis of the friend account names corresponding to each main communication account name in the fraud data to be analyzed, and the main communication account name and the target friend account name correspond to the same fraud group name; The friend account names corresponding to the main communication account name, excluding the target friend account name, form a first account name set; The account names of the group members of the group chat corresponding to the main communication account name, except the main communication account name and the target friend account name, are combined into a second account name set; An intersection of the first account name set and the second account name set is calculated to obtain identical account names included in the first account name set and the second account name set, and accounts corresponding to the identical account names are determined to be defrauded accounts.
5. The method according to any one of claims 1 to 4, characterized in that: The determining the fraud information of the victim account based on the chat data of the victim account in the fraud application comprises: Obtaining private chat data and group chat data of the defrauded account in the fraud application; Extracting multiple keywords from the private chat data and the group chat data, and determining weights corresponding to the multiple keywords based on the number of times the multiple keywords appear respectively; The multiple keywords are sorted in descending order of weight to obtain a target number of target keywords with top rankings, and the fraudulent information of the defrauded account is determined based on the target keywords.
6. The method according to claim 5, characterized in that The step of determining the fraud-related information of the fraud-related account based on the target keyword includes: Matching each communication record in the fraud data to be analyzed by using the target keyword; If the target keyword matches the target communication record, determining that the communication account corresponding to the target communication record is a fraudulent account, and outputting a list of fraudulent accounts; For each defrauded account in the defrauded account list, defrauded information of the defrauded account is extracted based on the contextual communication content in the communication record of the defrauded account.
7. A fraud data analysis device, characterized in that: The device comprises: A fraud data acquisition module, used to acquire fraud data to be analyzed; A first data classification module is used to classify the first fraud data with a main communication account name in the fraud data to be analyzed based on the main communication account name, and determine the first fraud data associated with the main communication account name as the same type of first fraud data to obtain a first fraud data classification result, wherein the main communication account name is the account name corresponding to the account that actively initiates the communication, and each type of first fraud data corresponds to a fraud group name; a second data classification module, configured to obtain wireless network names corresponding to all first fraudulent data in the first fraudulent data classification result, and update the first fraudulent data classification result based on the obtained wireless network name and the second fraudulent data to obtain a second fraudulent data classification result, wherein the same type of fraudulent data corresponds to the same wireless network name, and the second fraudulent data is the fraudulent data to be analyzed except the first fraudulent data; a fraudulent application determination module, configured to obtain, for each type of fraudulent data in the second fraudulent data classification result, an application installation list corresponding to all fraudulent data in each type of fraudulent data, count the number of installations of each application in the application installation list, and determine a social communication application with an installation number greater than a preset number as a fraudulent application; A fraudulent account determination module, for determining, for each type of fraudulent data in the second fraudulent data classification result, a fraudulent account based on the friend account names of each main communication account name in each type of fraudulent data and the group member account names of the group chat corresponding to each main communication account name; The fraudulent information determination module is used to determine the fraudulent information of the fraudulent account based on the chat data of the fraudulent account in the fraudulent application.
8. An electronic device, characterized in that: include: at least one processor; a memory for storing the at least one processor-executable instruction; The at least one processor is configured to execute the instructions to implement the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.