Electricity fraud susceptible population identification method and device, and storage medium
By using a risk identification rule engine with multi-source data and dynamic weights, combined with sentiment analysis and natural language processing, the system solves the problems of low efficiency and accuracy in identifying vulnerable groups for telecom fraud by financial institutions. It achieves automated and precise identification of vulnerable groups for telecom fraud, improving the efficiency and accuracy of early warning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technologies, financial institutions face challenges in accurately identifying vulnerable groups for telecom fraud due to low efficiency, high subjectivity, limited data dimensions, and lagging rule updates.
By acquiring multi-source data, including proprietary data from financial institutions, samples of victims of telecom fraud provided by external authorities, and behavioral data from third-party social media, multi-dimensional feature attributes are extracted to construct a risk identification rule set. Based on samples of victims of telecom fraud, weights are dynamically assigned to each rule to form a risk identification rule engine. This engine is used to score the accounts to be identified, and combined with sentiment analysis and natural language processing, a secondary risk analysis is conducted.
It enables automated and precise identification of vulnerable groups to telecom fraud, significantly improving early warning efficiency and accuracy, and helping financial institutions more effectively protect customer funds and reduce property losses.
Smart Images

Figure CN121834601A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial risk control, specifically to a method for identifying vulnerable groups to telecommunications fraud, a device for identifying vulnerable groups to telecommunications fraud, a processor, a machine-readable storage medium, and a computer program product. Background Technology
[0002] With the increasing prevalence of telecom fraud, accurately identifying and protecting vulnerable groups has become a crucial issue for financial institutions in risk control. Vulnerable groups are often more easily targeted by fraudsters due to a variety of factors, including age, socioeconomic status, psychological characteristics, and information literacy.
[0003] In existing technologies, financial institutions mainly use traditional identification methods based on human experience, such as customer identity verification and transaction behavior identification. These identification methods have limited data dimensions, lagging rule updates, low identification efficiency, and strong subjectivity. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, processor, storage medium, and machine-readable storage medium for identifying vulnerable groups to telecommunications fraud.
[0005] To achieve the above objectives, the first aspect of this application provides a method for identifying vulnerable groups to telecom fraud. The method includes: acquiring multi-source data, which includes at least: proprietary data from financial institutions, samples of telecom fraud victims provided by external authorized agencies, and social media behavior data provided by third parties; extracting multi-dimensional feature attributes based on the multi-source data; wherein the multi-dimensional feature attributes include: account attributes, customer attributes, transaction attributes, and social media behavior attributes; constructing a risk identification rule set based on the multi-dimensional feature attributes, and dynamically assigning weights to each rule based on the hit rate of the telecom fraud victim samples to each rule, forming a risk identification rule engine; scoring the telecom fraud victim samples using the risk identification rule engine, and using the average of all scores as a set threshold; determining the vulnerability score of the account to be identified using the risk identification rule engine, and determining whether the user corresponding to the account to be identified is a vulnerable group to telecom fraud based on the vulnerability score and the set threshold.
[0006] In this embodiment of the application, the step of extracting multi-dimensional feature attributes based on the multi-source data includes: extracting account information, customer information, and transaction records corresponding to the telecommunications fraud victim sample from proprietary data of financial institutions as telecommunications fraud victim account data; extracting victim sample behavior data corresponding to the telecommunications fraud victim sample from the social media behavior data; and extracting multi-dimensional feature attributes based on the telecommunications fraud victim sample, the telecommunications fraud victim account data, and the victim sample behavior data.
[0007] In this embodiment of the application, the step of dynamically allocating weights to each rule based on the hit rate of the fraud victim sample includes: counting the number of times each rule hits in the fraud victim sample; and allocating weights to each rule based on the proportion of the number of hits of each rule in the total number of hits.
[0008] In this embodiment of the application, the step of using the risk identification rule engine to determine the susceptibility score of the account to be identified includes: when the account to be identified matches a rule in the rule set, calculating the preset score corresponding to the matched rule; and determining the susceptibility score based on the weighted sum of the preset scores of all matched rules and their corresponding weights.
[0009] In this embodiment of the application, the method further includes: acquiring account interaction data generated by customers through mobile financial service applications or customer service hotlines; performing natural language processing on the account interaction data to extract emotional state features, fraud keyword features, and fraud speech pattern features; and constructing the risk identification rule set based on the emotional state features, fraud keyword features, and fraud speech pattern features.
[0010] In this embodiment of the application, the natural language processing includes at least one of the following: using a sentiment analysis model to determine the emotional state of the customer; using a keyword extraction model to extract keywords related to telecommunications fraud; and using a text classification model to identify patterns of telecommunications fraud tactics.
[0011] In this embodiment of the application, after determining the susceptibility score of the account to be identified using the risk identification rule engine, the method further includes: obtaining a training sample set, the training sample set including multi-dimensional feature attributes of people susceptible to and unsustainable in the field of telecommunications fraud; training an identification model using the training sample set; and, if the difference between the susceptibility score and the set threshold is less than the set difference, performing a secondary risk analysis on the account to be identified using the identification model to determine whether the user corresponding to the account to be identified is a person susceptible to telecommunications fraud.
[0012] A second aspect of this application provides a device for identifying vulnerable groups to telecom fraud. The device includes: a data acquisition module for acquiring multi-source data, including at least: proprietary data from financial institutions, samples of telecom fraud victims provided by external authorized agencies, and social media behavior data provided by third parties; an attribute extraction module for extracting multi-dimensional feature attributes based on the multi-source data, wherein the multi-dimensional feature attributes include: account attributes, customer attributes, transaction attributes, and social media behavior attributes; a rule engine formation module for constructing a risk identification rule set based on the multi-dimensional feature attributes, and dynamically assigning weights to each rule based on the hit rate of the telecom fraud victim samples to form a risk identification rule engine; a threshold determination module for scoring the telecom fraud victim samples using the risk identification rule engine, and using the average of all scores as a threshold; and an identification module for determining the vulnerability score of an account to be identified using the risk identification rule engine, and determining whether the user corresponding to the account to be identified is a vulnerable group to telecom fraud based on the vulnerability score and the threshold.
[0013] In this embodiment of the application, the step of extracting multi-dimensional feature attributes based on the multi-source data includes: extracting account information, customer information, and transaction records corresponding to the telecommunications fraud victim sample from proprietary data of financial institutions as telecommunications fraud victim account data; extracting victim sample behavior data corresponding to the telecommunications fraud victim sample from the social media behavior data; and extracting multi-dimensional feature attributes based on the telecommunications fraud victim sample, the telecommunications fraud victim account data, and the victim sample behavior data.
[0014] In this embodiment of the application, the step of dynamically allocating weights to each rule based on the hit rate of the fraud victim sample includes: counting the number of times each rule hits in the fraud victim sample; and allocating weights to each rule based on the proportion of the number of hits of each rule in the total number of hits.
[0015] In this embodiment of the application, the step of using the risk identification rule engine to determine the susceptibility score of the account to be identified includes: when the account to be identified matches a rule in the rule set, calculating the preset score corresponding to the matched rule; and determining the susceptibility score based on the weighted sum of the preset scores of all matched rules and their corresponding weights.
[0016] In this embodiment of the application, the device is further configured to: acquire account interaction data generated by customers through mobile financial service applications or customer service hotlines; perform natural language processing on the account interaction data to extract emotional state features, fraud keyword features, and fraud speech pattern features; and construct the risk identification rule set based on the emotional state features, fraud keyword features, and fraud speech pattern features.
[0017] In this embodiment of the application, the natural language processing includes at least one of the following: using a sentiment analysis model to determine the emotional state of the customer; using a keyword extraction model to extract keywords related to telecommunications fraud; and using a text classification model to identify patterns of telecommunications fraud tactics.
[0018] In this embodiment of the application, after determining the susceptibility score of the account to be identified using the risk identification rule engine, the device is further configured to: acquire a training sample set, the training sample set including multi-dimensional feature attributes of people susceptible to and unsustainable in the field of telecommunications fraud; train an identification model using the training sample set; and, if the difference between the susceptibility score and the set threshold is less than the set difference, perform a secondary risk analysis on the account to be identified using the identification model to determine whether the user corresponding to the account to be identified is a person susceptible to telecommunications fraud.
[0019] A third aspect of this application provides a processor configured to perform the method for identifying vulnerable groups to telecommunications fraud described above.
[0020] A fourth aspect of this application provides a machine-readable storage medium storing instructions that, when executed by a processor, configure the processor to perform the aforementioned method for identifying vulnerable groups to telecommunications fraud.
[0021] The fifth aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the method for identifying vulnerable groups to telecommunications fraud described above.
[0022] The technical solution provided in this application has at least the following technical effects: This application's method for identifying vulnerable groups to telecom fraud involves acquiring proprietary data from financial institutions, samples of telecom fraud victims from authorized external agencies, and social media behavior data obtained from third parties. This multi-source information is integrated to construct a more comprehensive and multi-dimensional user risk profile. Next, based on the multi-source data, multi-dimensional feature attributes such as accounts, customers, transactions, and social media behavior are extracted to construct a rule set. The hit rate of each rule is dynamically weighted according to the telecom fraud victim samples, enabling the rule engine to adapt to constantly changing telecom fraud methods, making the identification focus more prominent and targeted, ultimately forming a risk identification rule engine. Then, the rule engine is used to score the telecom fraud victim samples, and the average of all scores is used as a set threshold, making the risk judgment standard objective and scientific. Finally, the rule engine is used to determine the vulnerability score of the account to be identified. If the vulnerability score is greater than the set threshold, the user is identified as a vulnerable group to telecom fraud. This achieves automated and precise screening of vulnerable groups, significantly improving the efficiency and accuracy of early warnings, helping financial institutions more effectively protect customer funds and reduce financial losses. Therefore, the identification method for vulnerable groups to telecom fraud provided in this application can improve the accuracy of identification, enhance early warning efficiency, and effectively protect customer funds.
[0023] Other features and advantages of the embodiments of this application will be described in detail in the following detailed description section. Attached Figure Description
[0024] The accompanying drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the following detailed description to explain the embodiments of this application, but do not constitute a limitation on the embodiments of this application. In the drawings: Figure 1 A flowchart illustrating a method for identifying vulnerable groups to telecommunications fraud according to an embodiment of this application is shown in the schematic diagram. Figure 2 This illustration schematically shows a device for identifying vulnerable groups to telecommunications fraud according to an embodiment of this application; Figure 3 The diagram illustrates the internal structure of a computer device according to an embodiment of this application.
[0025] Explanation of reference numerals in the attached figures 200 - Identification device for vulnerable groups to telecom fraud; 201 - Data acquisition module; 202 - Attribute extraction module; 203 - Rule engine formation module; 204 - Threshold setting and determination module; 205 - Identification module; A01 - Processor; A02 - Network interface; A03 - Internal memory; A04 - Non-volatile storage medium; B01 - Operating system; B02 - Computer program. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for illustration and explanation of the embodiments of this application and are not intended to limit the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0027] It should be noted that if the embodiments of this application involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.
[0028] Furthermore, if the embodiments of this application involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0029] The acquisition, transmission, storage, use, and processing of data in this application comply with relevant national laws and regulations. Furthermore, it should be noted that existing industry solutions such as software, components, and models may be mentioned in the embodiments of this application. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.
[0030] Figure 1 This illustration schematically shows a flowchart of a method for identifying vulnerable groups to telecommunications fraud according to an embodiment of this application. Figure 1As shown in one embodiment of this application, a method for identifying vulnerable groups to telecom fraud is provided. The method includes: S101: acquiring multi-source data, which includes at least: proprietary data of financial institutions, samples of telecom fraud victims provided by external authorized agencies, and social media behavior data provided by third parties; S102: extracting multi-dimensional feature attributes based on the multi-source data; wherein, the multi-dimensional feature attributes include: account attributes, customer attributes, transaction attributes, and social media behavior attributes; S103: constructing a risk identification rule set based on the multi-dimensional feature attributes, and dynamically assigning weights to each rule based on the hit rate of each rule by the telecom fraud victim samples to form a risk identification rule engine; S104: scoring the telecom fraud victim samples using the risk identification rule engine, and using the average of all scores as a set threshold; S105: determining the vulnerability score of the account to be identified using the risk identification rule engine, and determining whether the user corresponding to the account to be identified is a vulnerable group to telecom fraud based on the vulnerability score and the set threshold.
[0031] Specifically, in this application's implementation, multi-source data is first collected for comprehensive analysis. This data includes at least three parts: first, authoritative samples of involved accounts and victims of telecom fraud provided by external authorized agencies; suspicious accounts that have transaction records with involved accounts within a set time period (e.g., three months) can also be classified as samples of victims of telecom fraud; second, account information, basic customer information, and all historical transaction records stored internally by financial institutions; and third, information on users' public behavior on social media platforms provided through compliant third-party channels. Next, based on this collected raw data, multi-dimensional characteristic attributes that comprehensively reflect potential risks are extracted. These attributes specifically cover four aspects: account attributes, customer attributes, transaction attributes, and social media behavior attributes. Account attributes include account number, account name, account status, account balance, account opening institution, account opening time, account opening location, account opening bank, and the location of the last transaction. Customer attributes include the user's age, gender, occupation, education, place of residence, and assets. Transaction attributes include the transaction date, transaction time, transaction amount, lending information, counterparty account number, transaction location, transaction channel, counterparty account number, and transaction summary for each transaction. Social media behavior attributes include the user's online activity, social network structure, sensitive information leakage indicators, and online shopping preferences.
[0032] Then, a rule set for risk identification is constructed using the extracted multi-dimensional feature attributes. The construction process includes: based on the statistical regularities of multi-dimensional features in the samples of victims of telecom fraud, the feature conditions with significant distinguishing power are solidified into specific rules. For example, specific rules are set such as the customer's age being greater than 60 or less than 30, the counterparty's account opening institution being in a different region from the account opening institution of the user's own account, and having more than five shopping records on niche platforms within a month. To make the rules more indicative, weights are dynamically assigned to each rule. The indicative strength of each rule is evaluated using known samples of victims of telecom fraud. The analysis found that the elderly (i.e., over 60 years old) and young people (i.e., 18 to 30 years old) are high-risk groups for telecom fraud, and they also have characteristics such as making multiple transfers to unfamiliar accounts in a short period of time or transferring abnormal amounts. Therefore, rules reflecting these behaviors have a higher hit rate and are given higher weights, while other rules with fewer hit rates have lower weights. All weights are summed up to one.
[0033] These rules, each with its own weight, are integrated to build an executable rule engine. If a victim of telecom fraud matches a rule, the default score for that rule (e.g., 100 points) is calculated. If the rule is not matched, no score is calculated. Finally, the scores of all matched rules are weighted and summed with their corresponding weights to obtain the total score for the account.
[0034] Next, the constructed rule engine is used to score each of the telecom fraud victim samples individually. The final score of each victim sample is then averaged, and this average is used as the threshold for subsequent risk assessment. Finally, for any account to be identified that needs risk screening, the rule engine is used to extract multi-dimensional feature attributes and match them with a rule set to calculate a susceptibility score. If the susceptibility score exceeds the set threshold, the user corresponding to that account is determined to be a high-risk individual susceptible to telecom fraud.
[0035] The method for identifying vulnerable groups to telecom fraud provided in this application integrates proprietary data from financial institutions, victim sample data from external agencies, and social media behavior data to construct a more comprehensive and multi-dimensional risk profile based on multi-source data. Dynamic weighting of rules based on real victim sample data allows the rule engine to adapt to constantly changing telecom fraud tactics, making the identification focus more prominent and targeted. Simultaneously, using the average score of victim samples as a threshold ensures that the risk assessment standard is objective and scientific. This enables automated and precise screening of vulnerable groups to telecom fraud, significantly improving the efficiency and accuracy of early warnings, and helping financial institutions more effectively protect customer funds and reduce financial losses.
[0036] In one implementation, the step of extracting multi-dimensional feature attributes based on the multi-source data includes: extracting account information, customer information, and transaction records corresponding to the telecommunications fraud victim samples from proprietary data of financial institutions, as telecommunications fraud victim account data; extracting victim sample behavior data corresponding to the telecommunications fraud victim samples from the social media behavior data; and extracting multi-dimensional feature attributes based on the telecommunications fraud victim samples, the telecommunications fraud victim account data, and the victim sample behavior data.
[0037] Specifically, in this implementation, account information, customer information, and transaction records corresponding to the fraud victim samples provided by external authorized agencies are first accurately extracted from proprietary data of financial institutions to form fraud victim account data. Account information specifically includes account number, account name, account status, account balance, account opening institution, account opening time and location; customer information specifically includes age, gender, occupation, education, place of residence, and asset status; transaction records specifically cover the date, time, amount, lending / borrowing information, counterparty account, transaction location, transaction channel, and transaction summary for each transaction. Victim sample behavior data corresponding to the victim sample data is extracted from social media behavior data provided by third parties, specifically including users' online activity, social network structure, indicators of sensitive information leakage, and online shopping preferences. Finally, based on the fraud victim account data, victim sample behavior data, and fraud victim samples, the extracted information is summarized and organized into four dimensions of feature attributes: account attributes, customer attributes, transaction attributes, and social media behavior attributes, laying a data foundation for subsequent risk identification.
[0038] The identification method for vulnerable groups to telecom fraud provided in this application can accurately depict the full picture of telecom fraud victims, enabling subsequent rule construction to be based on specific and quantifiable indicators, ensuring the comprehensiveness of the analysis perspective, effectively capturing risk signals that are easily overlooked by a single data source, thereby improving the accuracy and reliability of subsequent model identification.
[0039] In one embodiment, the step of dynamically assigning weights to each rule based on the hit rate of the fraud victim sample includes: counting the number of times each rule hits the fraud victim sample; and assigning weights to each rule based on the proportion of the number of hits of each rule in the total number of hits.
[0040] Specifically, in this embodiment, to assign a reasonable weight to each rule, firstly, each rule in the constructed risk identification rule set is applied one by one to all samples of victims of telecom fraud for matching, and the specific number of times each rule is hit is counted. For example, after analyzing victim data provided by external authorized agencies, it is found that the rule about customers being older than 60 or younger than 30 years old was hit 300 times in the victim samples, while the rule about shopping more than five times on niche platforms within a month was hit 50 times. Next, the proportion of each rule's hit count to the total number of hit counts of all rules is calculated, and this proportion is directly used as the weight of that rule. For example, if the total number of hit counts of all rules is 1, then the weight of age-related rules is 30%, and the weight of shopping-related rules is 5%. Weights are assigned to all rules in this way, ensuring that the sum of the weights of all rules is one.
[0041] The identification method for vulnerable groups to telecom fraud provided in this application can determine the importance of rules based on real telecom fraud victim cases, thereby improving the objectivity and accuracy of the judgment. When telecom fraud methods change, the number of times the rules reflecting the changes are hit in the victim sample increases, and the weight will also automatically increase accordingly. This allows for adaptive adjustment of the identification focus, making risk identification more focused on behaviors with high risk indicators, thereby improving the accuracy and efficiency of identification.
[0042] In one embodiment, determining the susceptibility score of an account to be identified using the risk identification rule engine includes: if the account to be identified matches a rule in the rule set, calculating the preset score corresponding to the matched rule; and determining the susceptibility score based on the weighted sum of the preset scores of all matched rules and their corresponding weights.
[0043] Specifically, in this embodiment, when determining the susceptibility score of an account to be identified, the multi-dimensional characteristic attributes of the account are first matched with each rule in the risk identification rule set. If a certain characteristic of the account matches a rule, a preset score, such as 100 points, is recorded for that rule; if it does not match, the rule is not scored. Then, the scores corresponding to all the matched rules are multiplied by their respective weights calculated from victim sample data, and these products are added together. The final sum is the total susceptibility score of the account to be identified.
[0044] The identification method for vulnerable groups to telecom fraud provided in this application can take into account the actual risk indication strength of each rule, so that the final vulnerability score can more accurately reflect the true risk level of the account, avoid the bias caused by simple counting, and provide a more reliable basis for risk ranking and accurate early warning.
[0045] In one embodiment, the method further includes: acquiring account interaction data generated by customers through mobile financial service applications or customer service hotlines; performing natural language processing on the account interaction data to extract sentiment features, fraud keyword features, and fraud speech pattern features; and constructing the risk identification rule set based on the sentiment features, fraud keyword features, and fraud speech pattern features.
[0046] In one embodiment, the natural language processing includes at least one of the following: using a sentiment analysis model to determine the customer's emotional state; using a keyword extraction model to extract fraudulent keywords; and using a text classification model to identify fraudulent script patterns.
[0047] Specifically, in this embodiment, the system proactively acquires customer chat logs from mobile financial service applications and text converted from voice messages during customer service calls, forming account interaction data. Next, natural language processing is performed on this account interaction data, specifically through the collaborative implementation of three models. First, a sentiment analysis model is used to perform in-depth analysis of the customer's dialogue text. By learning from a large amount of corpus containing emotions such as anxiety, panic, and urgency, the model can accurately determine the customer's emotional state during communication. For example, when a customer types text like "Has my account been frozen? What should I do? I'm very anxious" in a mobile financial service terminal chat, the sentiment analysis model will identify the panic and urgency contained within.
[0048] Secondly, a keyword extraction model is used to scan the text to identify and extract keywords highly related to telecom fraud, such as "safe account," "transfer verification," and "case number." When the dialogue includes phrases like "the other party asked me to transfer money to a designated safe account," the keyword extraction model accurately extracts the two key entities: the other party and the safe account.
[0049] Finally, a text classification model is used to perform pattern recognition on the context of the entire dialogue. By learning from a large number of known cases of telecom fraud, the text classification model can identify the complete telecom fraud dialogue process. For example, it can identify the typical dialogue pattern of first impersonating an authoritative figure to create a crisis, then offering solutions and urging urgent money transfers, thus determining that this belongs to the type of telecom fraud impersonating an authoritative institution, rather than ordinary business consultation. Finally, these newly extracted emotional states, telecom fraud keywords, and telecom fraud dialogue pattern features are transformed into specific recognition rules and added to the original risk identification rule set for more detailed risk assessment.
[0050] The method for identifying vulnerable groups to telecom fraud provided in this application can capture early warning signals before a scam occurs, identifying risks when customers are being scammed or have just come into contact with scammers. Furthermore, by combining emotional state and telecom fraud tactics, it can effectively distinguish between legitimate business inquiries and genuine scam scenarios, reducing false alarm rates and improving the accuracy of risk warnings.
[0051] In one embodiment, after determining the susceptibility score of the account to be identified using the risk identification rule engine, the method further includes: obtaining a training sample set, the training sample set including multi-dimensional feature attributes of people susceptible to and unsustainable in the fight against telecom fraud; training an identification model using the training sample set; and, if the difference between the susceptibility score and the set threshold is less than the set difference, performing a secondary risk analysis on the account to be identified using the identification model to determine whether the user corresponding to the account to be identified is a person susceptible to telecom fraud.
[0052] Specifically, in this embodiment, after calculating the susceptibility score of the account to be identified through the risk identification rule engine, a more refined secondary analysis is performed on accounts whose scores are very close to a preset threshold (e.g., an absolute difference of less than 5% or 10 points). First, a sample set is created, including individuals already identified as susceptible to telecom fraud and normal non-susceptible individuals, along with their corresponding dimensional feature attribute data. Then, an identification model (such as a gradient boosting decision tree, random forest, or neural network) is trained using the sample set to obtain the association patterns between susceptible and non-susceptible individuals across various feature dimensions. These association patterns are complex and not easily discovered through a single rule. Next, when the susceptibility score of the account to be identified is very close to the set threshold—for example, the difference between the score and the threshold is within a very small range—it indicates that the rule engine cannot make an accurate judgment. At this point, all multi-dimensional feature attributes of the account to be identified are input into the trained identification model. The identification model can comprehensively analyze all features, perform a comprehensive risk assessment, output the predicted probability that the account is susceptible to telecom fraud, determine whether this probability exceeds another preset probability threshold, and determine whether the user corresponding to the account belongs to the telecom fraud susceptible group.
[0053] The method for identifying vulnerable groups to telecom fraud provided in this application can discover features that are difficult to detect through rules through secondary analysis, accurately identify complex scenarios, and reduce false positives and false negatives. Furthermore, the model is applied only to accounts with scores close to the threshold, saving computational resources while ensuring efficient system operation.
[0054] In one embodiment, acquiring multi-source data further includes: acquiring device data and environmental data during customer operation; extracting multi-dimensional feature attributes further includes: extracting at least one of remote login features, proxy IP usage features, or device replacement features from the device data and environmental data.
[0055] Specifically, in this embodiment, device data and environmental data during customer operation are also collected. Device data includes information such as the device's operating system, browser model, and screen resolution. Environmental data mainly records the network environment during customer operation, especially the IP address and its corresponding geographical location. From this data, key risk characteristics can be extracted. For example, the characteristic of logging in from a different location determines whether the customer's login location is significantly different from their usual login city; for example, an account that has always logged in from Beijing suddenly has a login record in Hainan. The characteristic of using a proxy IP can detect whether the IP address used by the customer during login belongs to a proxy server, because telecom fraudsters often use proxy IPs to hide their real location. The characteristic of device change can determine whether the customer suddenly used a completely new, never-before-used device to log in to the account, especially when the new device is significantly different from the old device that the customer has been using for a long time. Then, a rule set is supplemented based on these characteristics. By analyzing device information and the login environment, abnormal behaviors such as logging in from a different location can be detected in a timely manner, thereby accurately identifying the risk of account theft before financial losses occur and achieving proactive security protection.
[0056] In one embodiment, the method further includes constructing a relationship network by using entity information from multi-source data, such as accounts, counterparties, devices, and IP addresses, as nodes, and their transfer or login relationships as edges. If account A has transferred money to account B, or account C frequently logs in using a certain mobile phone, a line is drawn connecting these nodes. In this way, all transfer and login relationships between accounts, devices, and IPs are clearly presented on the relationship network. Hidden risk information can be analyzed using this relationship network. For example, calculating how many lines separate ordinary accounts from known fraudulent accounts; if only one or two lines separate them, it indicates a close relationship and high risk. Another example is identifying the number of confirmed fraudulent accounts among adjacent accounts of an existing account. If 80% of the adjacent accounts are fraudulent, then this account is also highly likely to be an accomplice. This information analyzed from the relationship network can help accurately identify individual fraudsters and fraud groups.
[0057] Please refer to Figure 2This application provides a device 200 for identifying vulnerable groups to telecom fraud, comprising: a data acquisition module 201 for acquiring multi-source data, including at least: proprietary data of financial institutions, samples of telecom fraud victims provided by external authorized agencies, and social media behavior data provided by third parties; an attribute extraction module 202 for extracting multi-dimensional feature attributes based on the multi-source data, including: account attributes, customer attributes, transaction attributes, and social media behavior attributes; a rule engine formation module 203 for constructing a risk identification rule set based on the multi-dimensional feature attributes, and dynamically assigning weights to each rule based on the hit rate of the telecom fraud victim samples to form a risk identification rule engine; a threshold determination module 204 for scoring the telecom fraud victim samples using the risk identification rule engine, and using the average of all scores as a threshold; and an identification module 205 for determining the vulnerability score of the account to be identified using the risk identification rule engine, and determining whether the user corresponding to the account to be identified is a vulnerable group to telecom fraud based on the vulnerability score and the threshold.
[0058] In this embodiment of the application, the step of extracting multi-dimensional feature attributes based on the multi-source data includes: extracting account information, customer information, and transaction records corresponding to the telecommunications fraud victim samples from proprietary data of financial institutions as telecommunications fraud victim account data; extracting victim sample behavior data corresponding to the telecommunications fraud victim samples from social media behavior data; and extracting multi-dimensional feature attributes based on the telecommunications fraud victim account data, the victim sample behavior data, and the telecommunications fraud victim samples.
[0059] In this embodiment of the application, the step of dynamically allocating weights to each rule based on the hit rate of the fraud victim sample includes: counting the number of times each rule hits in the fraud victim sample; and allocating weights to each rule based on the proportion of the number of hits of each rule in the total number of hits.
[0060] In this embodiment of the application, the step of using the risk identification rule engine to determine the susceptibility score of the account to be identified includes: when the account to be identified matches a rule in the rule set, calculating the preset score corresponding to the matched rule; and determining the susceptibility score based on the weighted sum of the preset scores of all matched rules and their corresponding weights.
[0061] In this embodiment of the application, the device is further configured to: acquire account interaction data generated by customers through mobile financial service applications or customer service hotlines; perform natural language processing on the account interaction data to extract emotional state features, fraud keyword features, and fraud speech pattern features; and construct the risk identification rule set based on the emotional state features, fraud keyword features, and fraud speech pattern features.
[0062] In this embodiment of the application, the natural language processing includes at least one of the following: using a sentiment analysis model to determine the emotional state of the customer; using a keyword extraction model to extract keywords related to telecommunications fraud; and using a text classification model to identify patterns of telecommunications fraud tactics.
[0063] In this embodiment of the application, after determining the susceptibility score of the account to be identified using the risk identification rule engine, the device is further configured to: acquire a training sample set, the training sample set including multi-dimensional feature attributes of people susceptible to and unsustainable in the field of telecommunications fraud; train an identification model using the training sample set; and, if the difference between the susceptibility score and the set threshold is less than the set difference, perform a secondary risk analysis on the account to be identified using the identification model to determine whether the user corresponding to the account to be identified is a person susceptible to telecommunications fraud.
[0064] The device for identifying vulnerable groups to telecommunications fraud includes a processor and a memory. The aforementioned data acquisition module, attribute extraction module, rule engine formation module, threshold setting and determination module, and identification module are all stored as program units in the memory. The processor executes the aforementioned program modules stored in the memory to implement the corresponding functions.
[0065] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured, and methods for identifying vulnerable groups to telecom fraud can be implemented by adjusting kernel parameters.
[0066] The memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0067] This application provides a processor configured to perform the method for identifying vulnerable groups to telecommunications fraud described above.
[0068] This application provides a machine-readable storage medium storing instructions that, when executed by a processor, configure the processor to perform the method for identifying vulnerable groups to telecommunications fraud as described above.
[0069] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 3 As shown. Figure 3This schematic diagram illustrates the internal structure of a computer device according to an embodiment of the present application. The computer device includes a processor A01, a network interface A02, a memory (not shown), and a database (not shown) connected via a system bus. The processor A01 provides computing and control capabilities. The memory includes internal memory A03 and a non-volatile storage medium A04. The non-volatile storage medium A04 stores an operating system B01, a computer program B02, and a database (not shown). The internal memory A03 provides an environment for the operation of the operating system B01 and the computer program B02 stored in the non-volatile storage medium A04. The network interface A02 is used for communication with external terminals via a network connection. When the computer program B02 is executed by the processor A01, it implements a method for identifying individuals susceptible to telecommunications fraud.
[0070] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0071] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing an initialization program with the following method steps: acquiring multi-source data, the multi-source data including at least: proprietary data of financial institutions, samples of victims of telecommunications fraud provided by external authorized agencies, and social media behavior data provided by third parties; extracting multi-dimensional feature attributes based on the multi-source data; wherein, the multi-dimensional feature attributes include: account attributes, customer attributes, transaction attributes, and social media behavior attributes; constructing a risk identification rule set based on the multi-dimensional feature attributes, and dynamically assigning weights to each rule based on the hit rate of each rule by the samples of victims of telecommunications fraud, forming a risk identification rule engine; scoring the samples of victims of telecommunications fraud using the risk identification rule engine, and using the average of all scores as a set threshold; determining the susceptibility score of the account to be identified using the risk identification rule engine, and determining whether the user corresponding to the account to be identified is a person susceptible to telecommunications fraud based on the susceptibility score and the set threshold.
[0072] In one implementation, the step of extracting multi-dimensional feature attributes based on the multi-source data includes: extracting account information, customer information, and transaction records corresponding to the telecommunications fraud victim samples from proprietary data of financial institutions, as telecommunications fraud victim account data; extracting victim sample behavior data corresponding to the telecommunications fraud victim samples from the social media behavior data; and extracting multi-dimensional feature attributes based on the telecommunications fraud victim samples, the telecommunications fraud victim account data, and the victim sample behavior data.
[0073] In one embodiment, the step of dynamically assigning weights to each rule based on the hit rate of the fraud victim sample includes: counting the number of times each rule hits the fraud victim sample; and assigning weights to each rule based on the proportion of the number of hits of each rule in the total number of hits.
[0074] In one embodiment, determining the susceptibility score of an account to be identified using the risk identification rule engine includes: if the account to be identified matches a rule in the rule set, calculating the preset score corresponding to the matched rule; and determining the susceptibility score based on the weighted sum of the preset scores of all matched rules and their corresponding weights.
[0075] In one embodiment, the method further includes: acquiring account interaction data generated by customers through mobile financial service applications or customer service hotlines; performing natural language processing on the account interaction data to extract sentiment features, fraud keyword features, and fraud speech pattern features; and constructing the risk identification rule set based on the sentiment features, fraud keyword features, and fraud speech pattern features.
[0076] In one embodiment, the natural language processing includes at least one of the following: using a sentiment analysis model to determine the customer's emotional state; using a keyword extraction model to extract fraudulent keywords; and using a text classification model to identify fraudulent script patterns.
[0077] In one embodiment, after determining the susceptibility score of the account to be identified using the risk identification rule engine, the method further includes: obtaining a training sample set, the training sample set including multi-dimensional feature attributes of people susceptible to and unsustainable in the fight against telecom fraud; training an identification model using the training sample set; and, if the difference between the susceptibility score and the set threshold is less than the set difference, performing a secondary risk analysis on the account to be identified using the identification model to determine whether the user corresponding to the account to be identified is a person susceptible to telecom fraud.
[0078] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on at least one computer-usable storage medium (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0079] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0080] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0081] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0082] In a typical configuration, a computing device includes at least one processor (CPU), input / output interfaces, network interfaces, and memory.
[0083] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0084] Computer-readable media include both permanent and non-permanent, removable and non-removable media, which can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0085] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0086] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for identifying vulnerable groups to telecommunications fraud, characterized in that, The methods for identifying vulnerable groups to telecom fraud include: Acquire multi-source data, which includes at least: proprietary data of financial institutions, samples of victims of telecommunications fraud provided by external authorized agencies, and social media behavior data provided by third parties; Based on the multi-source data, multi-dimensional feature attributes are extracted; wherein, the multi-dimensional feature attributes include: account attributes, customer attributes, transaction attributes, and social media behavior attributes; A risk identification rule set is constructed based on the multi-dimensional feature attributes, and the weight of each rule is dynamically assigned based on the hit rate of the telecommunications fraud victim samples, thus forming a risk identification rule engine. The risk identification rule engine is used to score the samples of victims of telecommunications fraud, and the average of all scores is used as a set threshold. The risk identification rule engine is used to determine the susceptibility score of the account to be identified, and based on the susceptibility score and the set threshold, it is determined whether the user corresponding to the account to be identified is a vulnerable group to telecom fraud.
2. The method for identifying vulnerable groups to telecommunications fraud according to claim 1, characterized in that, The extraction of multi-dimensional feature attributes based on the multi-source data includes: Account information, customer information, and transaction records corresponding to the aforementioned samples of victims of telecommunications fraud are extracted from proprietary data of financial institutions and used as data on accounts of victims of telecommunications fraud. Extract victim sample behavior data corresponding to the telecommunications fraud victim samples from the social media behavior data; Based on the samples of victims of telecom fraud, the data of accounts of victims of telecom fraud, and the behavioral data of the victims, multi-dimensional feature attributes are extracted.
3. The method for identifying vulnerable groups to telecommunications fraud according to claim 1, characterized in that, The method of dynamically assigning weights to each rule based on the hit rate of the fraud victim samples includes: Count the number of times each rule is matched in the sample of victims of telecom fraud; Each rule is assigned a weight based on its percentage of hits in the total number of hits.
4. The method for identifying vulnerable groups to telecommunications fraud according to claim 1, characterized in that, The process of determining the vulnerability score of the account to be identified using the risk identification rule engine includes: If the account to be identified matches a rule in the rule set, calculate the preset score corresponding to the matched rule; The susceptibility score is determined by a weighted sum of the preset scores of all hit rules and their corresponding weights.
5. The method for identifying vulnerable groups to telecommunications fraud according to claim 1, characterized in that, The method further includes: Acquire customer account interaction data generated through mobile financial service applications or customer service hotlines; Natural language processing is performed on the account interaction data to extract sentiment features, fraud keyword features, and fraud rhetoric pattern features; The risk identification rule set is constructed based on emotional state features, telecom fraud keyword features, and telecom fraud rhetoric pattern features.
6. The method for identifying vulnerable groups to telecommunications fraud according to claim 5, characterized in that, The natural language processing includes at least one of the following: Use sentiment analysis models to determine customers' emotional state; Extracting keywords related to telecommunications fraud using a keyword extraction model; Use text classification models to identify patterns in telecom fraud tactics.
7. The method for identifying vulnerable groups to telecommunications fraud according to claim 1, characterized in that, After determining the vulnerability score of the account to be identified using the risk identification rule engine, the method further includes: Obtain a training sample set, which includes multi-dimensional feature attributes of people susceptible to and unsustainable in the fight against telecommunications fraud. The recognition model is trained using the training sample set; If the difference between the susceptibility score and the set threshold is less than the set difference, the identification model is used to perform a secondary risk analysis on the account to be identified to determine whether the user corresponding to the account to be identified is a vulnerable group to telecom fraud.
8. A device for identifying vulnerable groups to telecommunications fraud, characterized in that, The device for identifying vulnerable groups to telecom fraud includes: The data acquisition module is used to acquire multi-source data, which includes at least: proprietary data of financial institutions, samples of victims of telecommunications fraud provided by external authorized agencies, and social media behavior data provided by third parties; The attribute extraction module is used to extract multi-dimensional feature attributes based on the multi-source data; wherein, the multi-dimensional feature attributes include: account attributes, customer attributes, transaction attributes, and social media behavior attributes; The rule engine forming module is used to construct a risk identification rule set based on the multi-dimensional feature attributes, and dynamically assign weights to each rule based on the hit rate of the telecommunications fraud victim samples to form a risk identification rule engine. A threshold determination module is used to score the telecommunications fraud victim samples using the risk identification rule engine, and to use the average of all scores as the set threshold. The identification module is used to determine the susceptibility score of the account to be identified using the risk identification rule engine, and to determine whether the user corresponding to the account to be identified is a vulnerable group to telecom fraud based on the susceptibility score and the set threshold.
9. The device for identifying vulnerable groups to telecommunications fraud according to claim 8, characterized in that, The extraction of multi-dimensional feature attributes based on the multi-source data includes: Account information, customer information, and transaction records corresponding to the aforementioned samples of victims of telecommunications fraud are extracted from proprietary data of financial institutions and used as data on accounts of victims of telecommunications fraud. Extract victim sample behavior data corresponding to the telecommunications fraud victim samples from the social media behavior data; Based on the samples of victims of telecom fraud, the data of accounts of victims of telecom fraud, and the behavioral data of the victims, multi-dimensional feature attributes are extracted.
10. The device for identifying vulnerable groups to telecommunications fraud according to claim 8, characterized in that, The method of dynamically assigning weights to each rule based on the hit rate of the fraud victim samples includes: Count the number of times each rule is matched in the sample of victims of telecom fraud; Each rule is assigned a weight based on its percentage of hits in the total number of hits.
11. The device for identifying vulnerable groups to telecommunications fraud according to claim 8, characterized in that, The process of determining the vulnerability score of the account to be identified using the risk identification rule engine includes: If the account to be identified matches a rule in the rule set, calculate the preset score corresponding to the matched rule; The susceptibility score is determined by a weighted sum of the preset scores of all hit rules and their corresponding weights.
12. The device for identifying vulnerable groups to telecommunications fraud according to claim 8, characterized in that, The device is also used for: Acquire customer account interaction data generated through mobile financial service applications or customer service hotlines; Natural language processing is performed on the account interaction data to extract sentiment features, fraud keyword features, and fraud rhetoric pattern features; The risk identification rule set is constructed based on emotional state features, telecom fraud keyword features, and telecom fraud rhetoric pattern features.
13. The device for identifying vulnerable groups to telecommunications fraud according to claim 12, characterized in that, The natural language processing includes at least one of the following: Use sentiment analysis models to determine customers' emotional state; Extracting keywords related to telecommunications fraud using a keyword extraction model; Use text classification models to identify patterns in telecom fraud tactics.
14. The device for identifying vulnerable groups to telecommunications fraud according to claim 8, characterized in that, After determining the vulnerability score of the account to be identified using the risk identification rule engine, the device is further used to: Obtain a training sample set, which includes multi-dimensional feature attributes of people susceptible to and unsustainable in the fight against telecommunications fraud. The recognition model is trained using the training sample set; If the difference between the susceptibility score and the set threshold is less than the set difference, the identification model is used to perform a secondary risk analysis on the account to be identified to determine whether the user corresponding to the account to be identified is a vulnerable group to telecom fraud.
15. A processor, characterized in that, The method is configured to perform the identification method for vulnerable groups to telecommunications fraud as described in any one of claims 1 to 7.
16. A machine-readable storage medium storing instructions thereon, characterized in that, When executed by a processor, the instruction causes the processor to be configured to perform the method for identifying vulnerable groups to telecommunications fraud as described in any one of claims 1 to 7.
17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for identifying vulnerable groups to telecommunications fraud as described in any one of claims 1 to 7.