Method and device for identifying malicious user in marketing activity, medium and product
By analyzing the location and time information and communication behavior characteristics of operator data, combined with base station stay and travel information, potential malicious users are screened and verified, solving the problem of accurately identifying malicious users in online e-commerce platforms and improving identification efficiency and accuracy.
Patent Information
- Application Number
- CN202511266002.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-11-14
AI Technical Summary
In marketing activities on online e-commerce shopping platforms, it is difficult to accurately identify user groups who use unregistered phone cards and phone cards sold with personal information for malicious marketing, making it difficult to stop fraudulent activities. Existing technologies have low accuracy and efficiency in identification.
By analyzing users' carrier data, location and time information in both the location and time dimensions are obtained. Combined with base station stay and travel information, grid mapping is performed to screen out potential malicious users. Communication data is then used to verify their behavioral characteristics and identify malicious users.
It improved the accuracy of identifying malicious users in marketing campaigns, reduced false positives and false negatives, and effectively prevented the spread of malicious activities.
Smart Images

Figure CN120957142A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and specifically to a method, apparatus, medium, and product for identifying malicious users in marketing activities. Background Technology
[0002] In recent years, various marketing campaigns on online e-commerce platforms have frequently included promotions offering coupons or gifts for registering a mobile phone number. Some criminals use a SIM card pool device (a network communication hardware device that can insert multiple SIM cards for remote bulk internet access or phone calls) to insert multiple SIM cards to claim these coupons or gifts, resulting in extremely poor marketing effectiveness for the platforms. These actions constitute fraud. Assuming these SIM card pool devices are defined as malicious user groups, the existence of unregistered SIM cards ("black cards") and SIM cards obtained by users selling their personal information makes accurate identification of malicious user groups who specifically exploit promotional activities extremely difficult. The accuracy and efficiency are both low, making it difficult to prevent the spread of malicious activity. Summary of the Invention
[0003] At least one embodiment of this application provides a method, apparatus, medium, and product for identifying malicious users in marketing campaigns, addressing the problem of inaccurate identification of malicious users in marketing campaigns.
[0004] To solve the above-mentioned technical problems, this application is implemented as follows:
[0005] In a first aspect, embodiments of this application provide a method for identifying malicious users in marketing activities, including:
[0006] By utilizing the operator data of the user to be identified, location and time information related to the user's location dimension data and time dimension data can be obtained;
[0007] Based on the location and time information corresponding to each user, the first user with the potential for malicious intent is identified;
[0008] Based on the communication data between the first user and the first user, a malicious user is identified.
[0009] Optionally, using the operator data of the user to be identified, location and time information associated with the user's location dimension data and time dimension data can be obtained, including:
[0010] Using the operator data of the user to be identified, location dimension data including the user's base station stay information and time dimension data including the user's location travel information are obtained; the location dimension data is used to represent the user's daily base station location information; the time dimension data is used to represent the user's departure grid information and arrival grid information.
[0011] The base station location information is processed by grid mapping, and each base station location information is associated with a corresponding preset geographic grid to obtain the grid identifier corresponding to each base station, forming the base station-grid association data;
[0012] Using the unique identifier of the user to be identified as the association primary key, if the departure grid information and arrival grid information are consistent with the grid identifier of the base station location point in the corresponding time interval in the association data, then the location dimension data and the time dimension data are associated and bound to obtain fused data;
[0013] The fused data is sorted according to a preset time period to obtain the location and time information associated with the sorted location dimension data and time dimension data.
[0014] Optionally, based on the location and time information corresponding to each user, determine the first user with malicious intent, including:
[0015] Based on the location and time information of each user, user pairs with the same conditions are filtered out, and users in the user pairs who meet the preset latitude and longitude conditions are marked as potential accompanying users;
[0016] Determine the confidence level of the accompanying distance of the potential accompanying users;
[0017] The confidence level of the companion relationship of the user pair is calculated based on the product of the accompanying distance confidence level and the accompanying time period information corresponding to the potential accompanying user.
[0018] The confidence scores of the association relationships of all user pairs are sorted, and users whose association relationship confidence scores with known malicious users meet a preset threshold are identified as the first users with the possibility of malicious intent.
[0019] Optionally, determining the accompanying distance confidence of the potential accompanying user includes:
[0020] Determine the average accompaniment distance of potential accompaniment users across all time periods;
[0021] Determine the minimum of the average distance for each user across all time periods;
[0022] Determine the maximum value of the average distance for each user across all time periods;
[0023] The confidence level of the companion distance of the potential companion users is determined based on the average companion distance of the potential companion users, the minimum value of the average value, and the maximum value of the average value.
[0024] Optionally, based on the communication data between the first user and the first user, a malicious user is determined, including:
[0025] Based on the communication data of the first user, construct the feature vector of the first user;
[0026] Cluster all the first users and determine the first centroid after clustering; the first centroid represents the average value of all the first users in the j-th cluster on the k-th feature;
[0027] Using a preset clustering algorithm, the distance between each feature vector and all the first centroids is determined, and the feature vector is assigned to the cluster corresponding to the centroid with the smallest distance.
[0028] All feature vectors are assigned to clusters, and the centroid of each cluster is recalculated and updated.
[0029] Using the aforementioned preset clustering algorithm, the steps of feature vector allocation and centroid update are performed iteratively to determine the target feature vector until the cluster allocation no longer changes;
[0030] Based on the target feature vector, malicious users are identified.
[0031] Optionally, based on the target feature vector, determining the malicious user includes:
[0032] Based on the target feature vector, the features are sorted in sequence according to the application usage features, consumption features, and communication features. The first user who satisfies all three of the application usage features, consumption features, and communication features is identified as a malicious user.
[0033] Optionally, the feature vector of the first user includes at least one of the following:
[0034] Usage history of platform applications;
[0035] The amount paid by the first user;
[0036] The first user's spending amount;
[0037] The call duration of the first user;
[0038] The number of text messages sent to the first user.
[0039] Secondly, embodiments of this application provide an apparatus for identifying malicious users in marketing activities, comprising:
[0040] The first acquisition module is used to acquire location and time information associated with the user's location dimension data and time dimension data by utilizing the operator data of the user to be identified;
[0041] The first determination module is used to determine the first user with malicious intent based on the location and time information corresponding to each user.
[0042] The second determining module is used to determine the malicious user based on the communication data between the first user and the first user.
[0043] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.
[0044] Fourthly, embodiments of this application provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0045] Compared with existing technologies, the method, apparatus, medium, and product for identifying malicious users in marketing activities provided in this application utilize the operator data of the user to be identified to obtain location and time information associated with the user's location dimension data and time dimension data. This location and time information is determined by combining the user's location and time trajectory and communication behavior characteristics. Based on the location and time information corresponding to each user, a first user with malicious potential is identified. Based on the first user and the first user's communication data, a malicious user is identified. This solution identifies a suspicious first user by analyzing location and time information, and then analyzes the communication data of the first user to accurately identify malicious user groups in marketing activities, avoiding the problem of low accuracy of single-dimensional screening in existing technologies and improving the accuracy of identification. Attached Figure Description
[0046] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0047] Figure 1 One of the flowcharts illustrating a method for identifying malicious users in a marketing campaign, provided in an embodiment of this application;
[0048] Figure 2 A second flowchart illustrating a method for identifying malicious users in marketing activities provided in this application embodiment;
[0049] Figure 3 A schematic diagram of the structure of a device for identifying malicious users in marketing activities provided in an embodiment of this application. Detailed Implementation
[0050] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, without limiting the number of objects; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, "A or B" covers three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0051] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc., in the instruction sent. An indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.
[0052] As described in the background section, in the prior art, due to the existence of unregistered phone cards (or "black cards") circulating in the market and some users selling their personal information to others to obtain phone cards, it has become extremely difficult to accurately identify malicious user groups that specifically exploit promotional activities for profit. This makes it difficult to prevent the spread of malicious phenomena, and the accuracy rate of identifying malicious user groups is low. To solve the above problems, this application provides a method, device, medium, and product for identifying malicious users in marketing activities, which can reduce or avoid the occurrence of the above situations, improve the accuracy and efficiency of identifying malicious users, and improve the user experience.
[0053] Please refer to Figure 1 This application provides a method for identifying malicious users in marketing activities, comprising:
[0054] Step 11: Using the operator data of the user to be identified, obtain the location and time information associated with the user's location dimension data and time dimension data.
[0055] In this embodiment, the location and time information of the user to be identified is obtained from the operator data, such as base station access logs, network login records, and signal switching trajectories. This location and time information can be determined by fusing the user's physical location at different times with the user's trajectory data. For example, this location and time information can be represented as: User A is located in the coverage area of base station X at 8:00, and in the coverage area of base station Y at 9:30, etc. This application uses this location and time information to provide basic data for subsequent analysis of whether the user's activity patterns are abnormal.
[0056] Step 12: Based on the location and time information corresponding to each user, determine the first user with the potential for malicious intent;
[0057] In this application, by analyzing the location and time information obtained in step 11, a first user whose behavioral trajectory matches the characteristics of a malicious user is selected. This first user is considered a candidate group with the potential for malicious activity. For example, if the location and time of multiple users highly overlap, such as 10 SIM cards consistently connecting to the same base station within the same time period and having a fixed location for a long period, it may be SIM cards inserted in batches by a modem pool device. Modem pools are usually fixed in place, and multiple SIM cards share the same location. If the user's trajectory violates normal human activity patterns, such as switching base stations across 3 cities within 1 hour, far exceeding normal movement speed, it may be an abnormal SIM card with a virtual location or batch operation. If the user is in a non-residential area for a long period, such as the base station coverage area being a computer room or warehouse, rather than a residential or office area, and has no reasonable movement trajectory, it may be the location of a malicious device.
[0058] This application uses the location and time information corresponding to each user to initially narrow down the scope through trajectory anomalies, identify suspicious users that require further verification, and reduce the workload of subsequent analysis.
[0059] Step 13: Identify the malicious user based on the communication data between the first user and the first user.
[0060] It should be noted that the term "malicious" here may include malicious or illegal activities such as "fraud." For example, malicious users could be groups that use SIM card pools or black market cards to collect coupons in bulk.
[0061] In this embodiment of the application, for the first user selected in step 12, their communication data, such as call records, text message content, access to shopping applications (APP), etc., are analyzed to verify whether their behavior meets the characteristics of marketing fraud, and they are ultimately determined to be malicious users.
[0062] For example, the communication behavior is singular: only receiving verification code SMS messages, no normal voice calls, and frequent interactions with e-commerce platform servers, consistent with bulk registration and coupon redemption operations. Batch collaboration characteristics: multiple primary users have highly consistent communication targets, such as all interacting with the same batch of marketing activity interfaces, or the high frequency of keywords like "discount code" and "verification" in SMS content. SIM card pool association characteristics: multiple primary users' SIM cards were processed in the same batch, with dispersed locations but synchronized operation times, consistent with the characteristics of bulk resale of SIM cards followed by centralized use.
[0063] This application uses cross-validation of communication data and trajectory data to eliminate false positives of users with abnormal trajectories but legitimate identities, accurately identifying genuine malicious users. Through a two-step process—initial screening using location and time information in step 12, followed by communication data verification in step 13—combining the physical location characteristics and behavioral intent characteristics of malicious users, it achieves efficient identification of fraudulent users in marketing activities, avoiding misjudgments caused by relying on a single dimension (such as location or communication) and improving identification accuracy.
[0064] Optionally, step 11 above includes:
[0065] Using the operator data of the user to be identified, location dimension data including the user's base station stay information and time dimension data including the user's location travel information are obtained; the location dimension data is used to represent the user's daily base station location information; the time dimension data is used to represent the user's departure grid information and arrival grid information.
[0066] The base station location information is processed by grid mapping, and each base station location information is associated with a corresponding preset geographic grid to obtain the grid identifier corresponding to each base station, forming the base station-grid association data;
[0067] Using the unique identifier of the user to be identified as the association primary key, if the departure grid information and arrival grid information are consistent with the grid identifier of the base station location point in the corresponding time interval in the association data, then the location dimension data and the time dimension data are associated and bound to obtain fused data;
[0068] The fused data is sorted according to a preset time period to obtain the location and time information associated with the sorted location dimension data and time dimension data.
[0069] In this embodiment of the application, the operator data of the user to be identified is used to obtain location dimension data including user base station stay information, and time dimension data including user location travel information; the user base station stay information and user location travel information can be displayed in the form of a source table.
[0070] For example, user base station camping information is represented by a user base station camping day table, which can be shown in Table 1:
[0071] Table 1: User Base Station Stay Dates
[0072]
[0073]
[0074] For example, user base station stay information is represented by a user location travel calendar, which can be shown in Table 2:
[0075] Table 2: User Location Travel Schedule
[0076]
[0077]
[0078] In Tables 1 and 2, IMSI stands for International Mobile Subscriber Identity, a globally standardized number used to uniquely identify mobile users in mobile communication networks. In Table 2, od represents the origin and destination.
[0079] The location data in Tables 1 and 2 above are extracted, merged, and deduplicated. Further, basic data extraction can be performed. The extracted location data is processed, and is marked "1" when the source table is the user's mobile operator base station stay day table, and "2" when the source table is the user's location travel day table. The user's mobile operator base station stay day table records the user's daily mobile operator base station location information. Data extraction is performed on Table 1 above: Time periods with a stay duration > 0 are selected, such as mobile phone number (XXXX), latitude and longitude (A1, B1), district / county (YYYY), where the stay duration (minutes) from 9:00 to 10:00 is > 0. This forms a relationship group [XXXX,[9:00,10:00),YYYY,(A1,B1),1] between the user, time period, district / county, stay point, and identifier. The user's daily location trajectory information is obtained through multiple relationship groups.
[0080] The user location travel daily table records the user's departure grid and arrival grid location information. Data extraction is performed on Table 2 above: the departure grid data and arrival grid data for each mobile phone number are merged to form the location data for that mobile phone number. Entry or departure times that are not integer time intervals are recorded as integer time intervals. For example, mobile phone number (XXXX), departure grid entry time and departure grid departure time ([09:21:31, 09:40:51]), departure grid area (YYYY), and departure grid ID (A1, B1) form a relationship group [XXXX, [9:00, 10:00), YYYY, (A1, B1), 2] between user, time period, district / county, stop point, and identifier. The user's daily travel location trajectory information is obtained through multiple relationship groups.
[0081] It is important to note that grid entry or departure times may be non-integer time intervals, and need to be recorded separately according to whether they span integer time intervals. Specific rules are as follows: Rule 1: If both entry and departure times are within the same time interval, such as an entry time of 9:10 and a departure time of 9:25, then both fall within the 9:00-10:00 time interval. Generate a travel location point relationship group with an entry time of 9:10 and a departure time of 9:25, such as [XXXX,[9:00,10:00),YYYY,(A1,B1),2]. Rule 2: If entry and departure times span time intervals, such as an entry time of 9:10 and a departure time of 10:10, then they span two time intervals: 9:00-10:00 and 10:00-11:00. Generate two travel location point relationship groups, namely [XXXX,[9:00,10:00),YYYY,(A1,B1),2] and [XXXX,[10:00,11:00),YYYY,(A1,B1),2].
[0082] Optionally, this application also needs to fuse the user base station stay information and the user location travel information to obtain fused data. Further, refer to Tables 1 and 2 above for examples to illustrate data fusion and deduplication. Since the location data in the two basic location tables is duplicated, they need to be merged and deduplicated. When the time period and location (latitude and longitude accurate to four decimal places) of the same user number's stay location trajectory and travel location trajectory on the same day are the same, the time period and location latitude and longitude of the fused record are both taken from the stay location trajectory on the same day, and are marked as 0 to clearly indicate that it is a fused record.
[0083] For example: User's location trajectory for the day: [XXXX, [9:00, 10:00), YYYY, (A14567, B19867), 1]; User's travel location trajectory for the day: [XXXX, [9:00, 10:00), YYYY, (A14529, B19895), 2]; The merged record is:
[0084] [XXXX,[9:00,10:00),YYYY,(A14567,B19867),0].
[0085] Furthermore, the fused data is sorted according to a preset time period to obtain the sorted user location and time information; for example, a daily table of user location information can be output. In this application, each user record is sorted in ascending order by the time the location point entered, and a corresponding sequence number is generated. The daily table of user location information can be shown in Table 3.
[0086] Table 3: Daily User Location Information Table
[0087]
[0088] This application obtains sorted user location and time information through three steps: basic data extraction, data fusion and deduplication, and record sorting. This sorted user location and time information may include: the user's final user, time period, time period sorting, district / county, location point, and the relationship group between the identifier. The sorted user location and time information can be recorded in a user location information daily table, which provides data support for subsequent distance calculations.
[0089] Furthermore, by using the above-mentioned fusion records of user base station stay schedule and user location travel schedule, user trajectories are decomposed by time period to obtain sorted user location and time information. Then, by using the sorted user location and time information, the accompanying distance of users in the same period is calculated, a set distance threshold is analyzed, and the confidence level of user accompanying relationship is marked.
[0090] Optionally, step 12 above includes:
[0091] Based on the location and time information of each user, user pairs with the same conditions are filtered out, and users in the user pairs who meet the preset latitude and longitude conditions are marked as potential accompanying users;
[0092] Determine the confidence level of the accompanying distance of the potential accompanying users;
[0093] The confidence level of the companion relationship of the user pair is calculated based on the product of the accompanying distance confidence level and the accompanying time period information corresponding to the potential accompanying user.
[0094] The confidence scores of the association relationships of all user pairs are sorted, and users whose association relationship confidence scores with known malicious users meet a preset threshold are identified as the first users with the possibility of malicious intent.
[0095] In this embodiment, by analyzing the companion relationships between users, potential suspicious users highly associated with known malicious users are accurately identified. Companion relationships represent associations where locations are close and times overlap. First, user pairs with the same conditions are filtered from the location and time information of all users, such as user combinations with overlapping time dimensions. Specifically, this could mean that user A and user B both have location records between 8:00 and 10:00. Then, these user pairs are verified against preset latitude and longitude conditions. User pairs that meet the conditions are marked as potential companion users. Preset latitude and longitude conditions can be expressed as: latitude and longitude accurate to four decimal places being the same; or, for example, the difference in latitude and longitude between two users' locations being within a preset range, such as within 50 meters. Through the above steps, users whose locations are close within the same time period are initially identified. These users may belong to the same malicious device, such as a SIM card pool, where multiple cards have fixed and close physical locations.
[0096] Furthermore, the accompanying distance between potential accompanying user pairs within overlapping time periods is calculated to quantify the location association strength between users. The closer the distance, the more likely they are to belong to the same device or group physically.
[0097] Furthermore, the confidence score of the association relationship is calculated. The confidence score of the association relationship between user pairs is calculated by multiplying the association distance by the association time period. If two users are close (small association distance) and have a long association time (large time period information), the product result better reflects a strong association relationship. For example, an association distance of 30 meters × an association time of 2 hours = 60, indicating a high confidence score. Conversely, a large distance or short time results in a low confidence score. By using the confidence score of the association relationship, combining the degree of proximity and the length of time overlap, the closeness of the association between users is quantified, avoiding misjudgments based on a single dimension (only distance or only time).
[0098] After ranking all user pairs by their association confidence levels, the system focuses on user pairs formed with known malicious users. If the confidence level of a user's association with a known malicious user exceeds a preset threshold (e.g., the threshold is set to 50), and the confidence level of that user pair is 60%, then that user is marked as the first user with a high probability of malicious activity. By associating with known malicious users, potential malicious users are precisely located. For example, if user C is a known SIM card user, and user D has a high confidence level in their association with C, it indicates that D is likely also a malicious user from the same SIM card pool.
[0099] This application utilizes the core characteristics of malicious user groups (such as SIM card pools), which have fixed physical locations and multiple cards together. It quantifies user associations through the dual dimensions of proximity and time overlap, and combines the endorsement of known malicious users to efficiently screen out potential malicious users, providing a reliable candidate group for subsequent accurate judgment and reducing false positives and false negatives.
[0100] Further, determining the accompanying distance confidence of the potential accompanying user includes:
[0101] Determine the average accompaniment distance of potential accompaniment users across all time periods;
[0102] Determine the minimum of the average distance for each user across all time periods;
[0103] Determine the maximum value of the average distance for each user across all time periods;
[0104] The confidence level of the companion distance of the potential companion users is determined based on the average companion distance of the potential companion users, the minimum value of the average value, and the maximum value of the average value.
[0105] In this application, the potential accompanying users include user A and user B, and the accompanying distance confidence level is pd. AB , used to represent the standardized probability of distance. The calculation method is shown in the following formula:
[0106]
[0107] Among them, avg(d AB ) represents the average distance between users A and B over 24 time periods; in(avg(d XY )) represents the minimum of the average associated distance for each user group (X,Y) over 24 time periods; max(avg(d XY )) represents the maximum value of the average of the associated distances for each user group (X,Y) over 24 time periods.
[0108] In one specific embodiment, the accompanying distance calculation includes: selecting two users, such as user A and user B, who are on the same date, in the same district / county, and during the same time period (hourly interval), with their latitude and longitude (A14567, B19867) accurate to four decimal places (A1, 35.27). If the latitude and longitude of these two users are the same after four decimal places, it proves that the accompanying distance between these two users is less than 500 meters (variable parameter, preset distance for different scenarios, latitude and longitude accurate to different decimal places), and they are accompanying users. The accompanying distance between user A and user B is calculated using latitude and longitude and recorded in the accompanying user information daily table. The accompanying user information daily table is shown in Table 4 below.
[0109] Table 4: Accompanying User Information Daily Schedule
[0110] Serial Number Fields Field Chinese name Example 1 msisdn_A Mobile phone number A XXXX 2 msisdn_B Mobile phone number B 14937101984 3 follow_time_interval Accompanying period [9:00,10:00) 4 follow_distance Accompanying distance 480 5 countyid District / County Code YYYY 6 msisdn_A_location_point Location coordinates of mobile phone number A (A14567, B19867) 7 msisdn_B_location_point Location coordinates of mobile phone number B (A13967, B18713) 8 stat_day Daily Division 9 prov_id Provincial Division
[0111] Furthermore, this application utilizes a user companion algorithm to statistically analyze the proportion of companion time periods and companion distances between users A and B, defining companion relationships and calculating confidence levels. The confidence level of a companion relationship is defined as follows: the reliability of a companion relationship is affected by two parameters: the number of companion time periods covered and the companion distance. The more companion time periods and the closer the companion distance, the more reliable the companion relationship. Therefore, the method for calculating the confidence level of a companion relationship is defined as follows: α AB =pt AB *pd AB ;
[0112] Wherein, the parameter α is defined AB This represents the confidence level of the association between users A and B; pt AB The confidence level for the accompanying period represents the percentage of time periods covered by accompanying records within a specific time range of A and B. It is calculated by dividing the number of accompanying periods on a given day by 24 (there are 24 time periods per day). Then, using the formula above: Calculate the confidence level pd of the associated distance AB .
[0113] Finally, the confidence scores of all user combinations are sorted from highest to lowest as the confidence score ranking, and recorded in the daily table of associated relationship results. The daily table of associated relationship results can be represented as Table 5.
[0114] Table 5: Daily Schedule of Accompanying Relationship Results
[0115] Serial Number Fields Field Chinese name Field Description 1 Msisdn_A Mobile phone number A 2 Msisdn_B Mobile phone number B 3 pt Accompanying time period confidence 4 pd Accompanying distance confidence 5 p Confidence of companion relationship In the formula, α has a range of [0,1]. 6 p_order Confidence ranking 7 countyid District / County Code YYYY 8 stat_day Daily Division 9 prov_id Provincial Division
[0116] Next, a probability analysis of the accompanying outcome is performed to determine the first user. The probability analysis of the accompanying outcome includes: (1) If user A and user B are in the same location for 24 time periods on a certain day and do not move, the confidence of their accompanying relationship is 1; (2) If user A and user B move within 24 time periods on a certain day, the confidence of the accompanying relationship approaches 1 as the number of accompanying time periods approaches 24 and the accompanying distance decreases; when the number of accompanying time periods is 24 and the accompanying distance is 0, the confidence of the accompanying relationship is 1; (3) If user A and user B move within 24 time periods on a certain day, the confidence of the accompanying relationship approaches 0 as the number of accompanying time periods approaches 0 and the accompanying distance increases.
[0117] When the confidence level of the association relationship is 1: Highly correlated potential malicious users are directly identified. A scenario with a confidence level of 1 (user A and user B's locations completely overlap and times are synchronized across 24 time periods) represents an extreme manifestation of a strong association relationship, corresponding to two typical cases: Case 1: Users A and B are in the same location throughout the day (24 time periods) and do not move (e.g., multiple SIM cards inserted into a SIM card pool device, physically fixed in a server room or warehouse, with no movement trajectory); Case 2: Users A and B move throughout the day but are completely synchronized (in the same location for all 24 time periods, with a distance of 0, such as multiple SIM cards inserted into the same phone simultaneously, or devices from the same malicious group operating synchronously). This characteristic of completely overlapping locations and synchronized times highly matches the core behavior of malicious user groups (using devices like SIM card pools for batch operations and physical location binding).
[0118] Therefore, the logic for determining the first user is as follows: if at least one party in a user pair is a known malicious user (such as a confirmed SIM card user), then the other party, due to the confidence level of the accompanying relationship = 1, indicates that the two are very likely to belong to the same malicious device or group (physical binding, collaborative operation), and is directly marked as the first user with the possibility of malicious intent; even if there are no known malicious users, such a group of users who are 100% accompanied throughout the day, such as 10 users with a confidence level of 1 for each other, will also be included in the scope of the first user because their behavior does not conform to the characteristics of normal users. Normal users cannot have completely overlapping locations for 24 hours, and they will wait for subsequent communication data verification.
[0119] II. When the confidence level of the association relationship is 0: Excluding the possibility of malicious association. In scenarios where the confidence level of the association relationship is 0, users A and B have almost no time overlap and are extremely far apart, representing an extreme case of no association: their location records have almost no time overlap throughout the 24 time periods of a day, the time overlap is close to 0, and their locations are extremely far apart, such as user A in Shanghai and user B in Guangzhou, with no location intersection throughout the day. This characteristic of zero association and long distance is completely inconsistent with the behavioral logic of malicious user groups. Malicious users need to operate through the same device or in a group, which inevitably leads to proximity in location and time overlap. Therefore, the logic for determining the first user is: even if one party in the user pair is a known malicious user, the other party, due to an association confidence level of 0, indicates that there is no physical connection or possibility of collaborative operation between the two, and their malicious possibility is directly excluded, and they are not marked as the first user; the location and time trajectories of such users are independent and unrelated, which is more in line with the movement patterns of normal users, such as independent users in different regions, and do not need to be included in the subsequent verification scope.
[0120] Optionally, step 13 above includes:
[0121] Based on the communication data of the first user, construct the feature vector of the first user;
[0122] Cluster all the first users and determine the first centroid after clustering; the first centroid represents the average value of all the first users in the j-th cluster on the k-th feature;
[0123] Using a preset clustering algorithm, the distance between each feature vector and all the first centroids is determined, and the feature vector is assigned to the cluster corresponding to the centroid with the smallest distance.
[0124] All feature vectors are assigned to clusters, and the centroid of each cluster is recalculated and updated.
[0125] Using the aforementioned preset clustering algorithm, the steps of feature vector allocation and centroid update are performed iteratively to determine the target feature vector until the cluster allocation no longer changes;
[0126] Based on the target feature vector, malicious users are identified.
[0127] It should be noted that the communication data of the first user can include the ratio of the number of e-commerce shopping platform apps used in the past three months to the total number of apps used in the past three months, the amount of money paid in the past three months, the amount of money spent in the past three months, the duration of calls in the past three months, and the number of text messages sent in the past three months. It should also be noted that "past three months" is just an example; other values can be used, and this application does not impose any restrictions.
[0128] It should also be noted that verification can be performed on all first users to obtain their communication data; alternatively, verification can be performed on a subset of first users. For example, a user group with a confidence score within 200,000 for a specific date and district code (200,000 is just an example, other values are possible) can be selected to obtain their communication data, including: the ratio of the number of e-commerce shopping platform apps used in the past 3 months to the total number of apps used in the past 3 months, the amount of money paid in the past 3 months, the amount of money spent in the past 3 months, the duration of calls in the past 3 months, and the number of text messages sent in the past 3 months.
[0129] Optionally, a feature vector of the first user is constructed based on the communication data of the first user. The feature vector of the first user includes at least one of the following:
[0130] Usage history of platform applications;
[0131] The amount paid by the first user;
[0132] The first user's spending amount;
[0133] The call duration of the first user;
[0134] The number of text messages sent to the first user.
[0135] In this embodiment of the application, a feature vector of the first user is constructed based on the communication data of the first user. The mobile phone number corresponding to each first user can be represented as a feature vector x. i Where i represents the index of the number, x i The number contains five characteristics: x i =(x i1 ,x i2 ,x i3 ,x i4 ,x i5 ); where x i1 This represents usage records for platform applications, such as the ratio of the number of e-commerce shopping platform apps used in the past three months to the total number of apps used in the past three months; x i2 This represents the payment amount of the first user, such as the payment amount for the past 3 months; x i3 This represents the first user's spending amount, such as spending amount over the past 3 months; x i4 This indicates the call duration of the first user, such as the call duration over the past 3 months; x i5 This indicates the number of text messages sent to the first user, such as the number of text messages sent in the past 3 months.
[0136] Furthermore, all the first users requiring verification are clustered, and a first centroid is determined after clustering; the first centroid represents the average value of all first users in the j-th cluster on the k-th feature; the first centroid μ j It is the center point of the j-th cluster, and also an eigenvector, represented as: μ j =(μ j1 ,μ j2 ,μ j3 ,μ j4 ,μ j5 ), where each μ jk It is the average value of all numbers in the j-th cluster on the k-th feature. First centroid selection: The base user group is randomly divided into 10 clusters (10 is just an example), and the first centroid μ is selected using the method described above. j .
[0137] The preset clustering algorithm in this application can be K-means clustering algorithm, which determines the distance between each feature vector and all the first centroids, and assigns the feature vector to the cluster corresponding to the centroid with the smallest distance. Specifically, in K-means clustering, Euclidean distance is used to measure the feature vector x. i With the center of mass μ j The distance between them is expressed as:
[0138]
[0139] This distance measures the feature vector x iWith the center of mass μ j Straight-line distance in all feature spaces.
[0140] All feature vectors are assigned to clusters, and the centroid of each cluster is recalculated and updated, i.e., for each feature vector x... i (That is, for each number), calculate its distance to each centroid and assign it to the cluster corresponding to the centroid with the smallest distance. This means that for each feature vector x i Find the centroid μ that has the smallest distance from its eigenvector. j And it is assigned to the j-th cluster.
[0141] Once all feature vectors (i.e. all numbers) have been assigned to clusters, recalculate the centroid of each cluster so that it is the mean of all points in that cluster: Among them, |C j | represents the number of eigenvectors in the j-th cluster, and the summation is performed on all eigenvectors belonging to the j-th cluster. This formula essentially calculates the average of all numbers in each cluster for each feature.
[0142] The K-means clustering algorithm iteratively performs eigenvector assignment and centroid update steps to determine the target eigenvector until the cluster assignment no longer changes. The key to the K-means algorithm is the iterative execution of eigenvector assignment and centroid update steps until the cluster assignment no longer changes. Each eigenvector x... i Users are assigned to a cluster, where the centroid of the cluster represents the characteristics of that cluster, thus forming a target feature vector. Finally, malicious users are identified based on this target feature vector.
[0143] Optionally, based on the target feature vector, determining the malicious user includes:
[0144] Based on the target feature vector, the features are sorted in sequence according to the application usage features, consumption features, and communication features. The first user who satisfies all three of the application usage features, consumption features, and communication features is identified as a malicious user.
[0145] In this embodiment, the target feature vector includes multiple features, which are ordered sequentially according to application usage features, consumption features, and communication features. The first user who satisfies all three of the application usage features, consumption features, and communication features is determined to be a malicious user. Based on the above target feature vector, the importance of the five features that the target feature vector can contain, ordered from highest to lowest, is as follows:
[0146] Application usage characteristics: the ratio of the number of e-commerce shopping platform apps used in the past 3 months to the total number of apps used in the past 3 months; consumption characteristics: the amount of bills paid and the amount of consumption in the past 3 months; communication characteristics: the call duration and the number of text messages sent in the past 3 months.
[0147] Explanation of feature sorting: The ratio of the number of times e-commerce shopping platform apps (such as Taobao) were used in the past 3 months to the total number of apps used in the past 3 months reflects the types of apps users frequently use. If a user is malicious, they will log in to e-commerce shopping platform apps to participate in marketing and promotional activities, so their ratio will be abnormally high. This cluster of users has a very high probability of being a malicious user group. Let's assume the selected clusters are C1, C2, C3, C4, and C5.
[0148] Based on clusters C1, C2, C3, C4, and C5, we continue to observe the characteristics of payment amounts and consumption amounts over the past three months for each cluster. Malicious users log into e-commerce shopping platforms online. To reduce fraud costs, they only maintain basic phone card services, hence their payment amounts over the past three months will be very low. Consequently, their consumption amounts over the past three months will also be low. These two indicators reflect user spending patterns and, consequently, overall user activity. Such users are highly likely to be malicious users. Let's assume the selected clusters are C1, C2, and C3.
[0149] Finally, based on clusters C1, C2, and C3, we observed the characteristics of call duration and SMS frequency over the past three months.
[0150] These two features reflect information about the user's social circle. If the call duration and text message frequency are low in the past three months, it indicates that the user's social circle is small, which is consistent with the characteristics of malicious user groups who spend more time online and less time communicating, further confirming that the identified cluster is a malicious user group.
[0151] For example, in the process of identifying malicious user groups, suppose there are 10 clusters C1, C2, C3, C4, C5, C6, C7, C8, C9, and C10. Observe the feature vector of their centroids. This feature vector reflects the average value of each feature of the cluster. Based on this value, the overall characteristics of the cluster can be determined. Based on the centroid feature μ... j1 (The average ratio of the number of e-commerce shopping platform apps used in the past 3 months to the total number of apps used in the past 3 months) was used to select clusters with a ratio greater than or equal to 0.8, resulting in clusters C1, C2, C3, C4, and C5. The centroid characteristics μ of clusters C1, C2, C3, C4, and C5 were then observed. j2 (Average payment amount over the past 3 months), μ j3 (Average spending over the past 3 months), filtered to μ j2 <=20 and μ j3 Clusters with values <= 5 are obtained as clusters C1, C2, and C3. The centroid characteristics μ of clusters C1, C2, and C3 are observed. j4(Average call duration over the past 3 months), μ j5 (Average number of text messages in the past 3 months), filtered to μ j4 <=5 minutes and μ j5 Clusters with a power of 3 or less are identified as cluster C1 (which may contain one or more clusters). At this point, the user group in cluster C1 is output as a malicious user group.
[0152] Reference Figure 2 As shown in the embodiments of this application, a method for identifying malicious users in marketing activities is also provided, including:
[0153] Step 21: Obtain mobile operator data, such as user location information and user communication data, for users participating in marketing activities; Step 22: Based on user location information, output users with similar travel trajectories as the basic user group using a user accompaniment algorithm; Step 23: Using the basic user group output in Step 22 and the communication information of this user group, output the fraudulent user group using a K-means clustering algorithm.
[0154] In summary, this application proposes fusing users' base station dwell time tables and location travel time tables to obtain users' location and time information; calculating the confidence level of the association relationship between user pairs using the location and time information of each user; and, based on the confidence level of the association relationship between user pairs, initially screening out users with malicious potential from those participating in marketing activities; for the initially screened users, determining malicious users based on their communication data (including: the ratio of the number of e-commerce shopping platform apps used to the total number of apps used within a predetermined time period, the payment amount within the predetermined time period, the consumption amount within the predetermined time period, the call duration within the predetermined time period, and the number of SMS messages within the predetermined time period). This application also proposes a method for calculating the confidence level of the association relationship between user pairs.
[0155] This application uses the confidence level of the companion relationship between user pairs for initial screening, which initially identifies users with companion relationships, greatly narrowing the user range and improving the overall screening efficiency. Then, it collects several communication data of the companion user groups identified in the initial screening, and then judges malicious user groups for the second step of clustering screening. It expands multiple screening dimensions, avoids large errors caused by a single dimension, and improves the accuracy of overall malicious user group identification.
[0156] When initially screening accompanying users, the user's base station stay schedule and location travel schedule are fused to obtain the user's time and location information. The advantage is that the base station stay data and location travel data complement each other, reducing the possibility of data loss, improving data completeness, and thus improving the accuracy of identification.
[0157] The method for calculating the confidence of the companion relationship between user pairs in the initial screening is based on two dimensions of data: time period and latitude and longitude. Multi-dimensional screening improves accuracy. The companion relationship confidence algorithm is used to identify the companion user group, which improves efficiency and accuracy.
[0158] The second screening step of this application utilizes the K-means clustering algorithm. Its advantages lie in the fact that the algorithm model is automatically calculated, which can quickly classify users and improve computational efficiency; and it can intuitively identify abnormal user groups, which improves convenience and accuracy.
[0159] The various methods described above are based on embodiments of this application. Apparatus for implementing the above methods will now be provided.
[0160] Please refer to Figure 3 This application also provides an apparatus for identifying malicious users in marketing activities, comprising:
[0161] The first acquisition module 31 is used to acquire location and time information associated with the user's location dimension data and time dimension data using the operator data of the user to be identified;
[0162] The first determination module 32 is used to determine the first user with malicious intent based on the location and time information corresponding to each user.
[0163] The second determining module 33 is used to determine a malicious user based on the communication data between the first user and the first user.
[0164] Optionally, the first acquisition module 31 described above includes:
[0165] The first acquisition unit is used to acquire location dimension data, including user base station stay information, and time dimension data, including user location travel information, using the operator data of the user to be identified; the location dimension data is used to represent the user's daily base station location information; the time dimension data is used to represent the user's departure grid information and arrival grid information.
[0166] The first processing unit is used to perform grid mapping processing on the base station location information, associate each base station location information with the corresponding preset geographic grid, obtain the grid identifier corresponding to each base station, and form base station-grid association data;
[0167] The second acquisition unit is used to associate and bind the location dimension data and the time dimension data with the unique identifier of the user to be identified as the association primary key. If the departure grid information and the arrival grid information are consistent with the grid identifier of the base station location point in the corresponding time interval in the association data, the fused data is obtained.
[0168] The third acquisition unit is used to sort the fused data according to a preset time period and acquire the location and time information associated with the sorted location dimension data and time dimension data.
[0169] Optionally, the first determining module 32 described above includes:
[0170] The second processing unit is used to filter out user pairs with the same conditions based on the location and time information corresponding to each user, and mark users in the user pairs who meet the preset latitude and longitude conditions as potential accompanying users.
[0171] The first determining unit is used to determine the confidence level of the accompanying distance of the potential accompanying user;
[0172] The third processing unit is used to calculate the confidence of the companion relationship of the user pair based on the product of the companion distance confidence and the companion time period information corresponding to the potential companion user;
[0173] The fourth processing unit is used to sort the confidence of the association relationship of all user pairs, and to determine the first user with the possibility of being malicious by identifying the user whose confidence of the association relationship with the known malicious user meets a preset threshold.
[0174] Optionally, the first determining unit described above is specifically used for:
[0175] Determine the average accompaniment distance of potential accompaniment users across all time periods;
[0176] Determine the minimum of the average distance for each user across all time periods;
[0177] Determine the maximum value of the average distance for each user across all time periods;
[0178] The confidence level of the companion distance of the potential companion users is determined based on the average companion distance of the potential companion users, the minimum value of the average value, and the maximum value of the average value.
[0179] Optionally, the second determining module 33 described above includes:
[0180] The fifth processing unit is used to construct the feature vector of the first user based on the communication data of the first user;
[0181] The third determining unit is used to cluster all the first users and determine the first centroid after clustering; the first centroid represents the average value of all the first users in the j-th cluster on the k-th feature;
[0182] The fourth determining unit is used to determine the distance between each feature vector and all the first centroids using a preset clustering algorithm, and to assign the feature vector to the cluster corresponding to the centroid with the smallest distance;
[0183] The sixth processing unit is used to assign all feature vectors to clusters, recalculate and update the centroid of each cluster;
[0184] The seventh processing unit is used to iteratively execute the steps of feature vector allocation and centroid update using the preset clustering algorithm to determine the target feature vector until the cluster allocation no longer changes;
[0185] The fifth determining unit is used to determine the malicious user based on the target feature vector.
[0186] Optionally, the fifth determining unit mentioned above is specifically used for:
[0187] Based on the target feature vector, the features are sorted in sequence according to the application usage features, consumption features, and communication features. The first user who satisfies all three of the application usage features, consumption features, and communication features is identified as a malicious user.
[0188] Optionally, the feature vector of the first user includes at least one of the following:
[0189] Usage history of platform applications;
[0190] The amount paid by the first user;
[0191] The first user's spending amount;
[0192] The call duration of the first user;
[0193] The number of text messages sent to the first user.
[0194] It should be noted that the device in this embodiment is a device corresponding to the method described above for identifying malicious users in marketing activities. The implementation methods in the above embodiments are all applicable to the embodiments of this device and can achieve the same technical effect. The device provided in this application embodiment can implement all the method steps implemented in the above method embodiments and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiments and the beneficial effects will not be described in detail.
[0195] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described method embodiment for identifying malicious users in marketing activities, and achieves the same technical effect. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0196] This application also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the various processes of the above-described method embodiment for identifying malicious users in marketing activities and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0197] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0198] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0199] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for identifying malicious users in marketing campaigns, characterized in that, include: By utilizing the operator data of the user to be identified, location and time information related to the user's location dimension data and time dimension data can be obtained; Based on the location and time information corresponding to each user, the first user with the potential for malicious intent is identified; Based on the communication data between the first user and the first user, a malicious user is identified.
2. The method according to claim 1, characterized in that, Using the carrier data of the user to be identified, obtain location and time information associated with the user's location and time dimensions, including: Using the operator data of the user to be identified, location dimension data including the user's base station stay information and time dimension data including the user's location travel information are obtained; the location dimension data is used to represent the user's daily base station location information; the time dimension data is used to represent the user's departure grid information and arrival grid information. The base station location information is processed by grid mapping, and each base station location information is associated with a corresponding preset geographic grid to obtain the grid identifier corresponding to each base station, forming the base station-grid association data; Using the unique identifier of the user to be identified as the association primary key, if the departure grid information and arrival grid information are consistent with the grid identifier of the base station location point in the corresponding time interval in the association data, then the location dimension data and the time dimension data are associated and bound to obtain fused data; The fused data is sorted according to a preset time period to obtain the location and time information associated with the sorted location dimension data and time dimension data.
3. The method according to claim 1, characterized in that, Based on the location and time information of each user, the first user with malicious intent was identified, including: Based on the location and time information of each user, user pairs with the same conditions are filtered out, and users in the user pairs who meet the preset latitude and longitude conditions are marked as potential accompanying users; Determine the confidence level of the accompanying distance of the potential accompanying users; The confidence level of the companion relationship of the user pair is calculated based on the product of the accompanying distance confidence level and the accompanying time period information corresponding to the potential accompanying user. The confidence scores of the association relationships of all user pairs are sorted, and users whose association relationship confidence scores with known malicious users meet a preset threshold are identified as the first users with the possibility of malicious intent.
4. The method according to claim 3, characterized in that, Determining the accompanying distance confidence of the potential accompanying user includes: Determine the average accompaniment distance of potential accompaniment users across all time periods; Determine the minimum of the average distance for each user across all time periods; Determine the maximum value of the average distance for each user across all time periods; The confidence level of the companion distance of the potential companion users is determined based on the average companion distance of the potential companion users, the minimum value of the average value, and the maximum value of the average value.
5. The method according to claim 1, characterized in that, Based on the communication data between the first user and the first user, a malicious user is identified, including: Based on the communication data of the first user, construct the feature vector of the first user; Cluster all the first users and determine the first centroid after clustering; the first centroid represents the average value of all the first users in the j-th cluster on the k-th feature; Using a preset clustering algorithm, the distance between each feature vector and all the first centroids is determined, and the feature vector is assigned to the cluster corresponding to the centroid with the smallest distance. All feature vectors are assigned to clusters, and the centroid of each cluster is recalculated and updated. Using the aforementioned preset clustering algorithm, the steps of feature vector allocation and centroid update are performed iteratively to determine the target feature vector until the cluster allocation no longer changes; Based on the target feature vector, malicious users are identified.
6. The method according to claim 5, characterized in that, Based on the target feature vector, malicious users are identified, including: Based on the target feature vector, the features are sorted in sequence according to the application usage features, consumption features, and communication features. The first user who satisfies all three of the application usage features, consumption features, and communication features is identified as a malicious user.
7. The method according to claim 5, characterized in that, The feature vector of the first user includes at least one of the following: Usage history of platform applications; The amount paid by the first user; The first user's spending amount; The call duration of the first user; The number of text messages sent to the first user.
8. A device for identifying malicious users in a marketing campaign, characterized in that, include: The first acquisition module is used to acquire location and time information associated with the user's location dimension data and time dimension data by utilizing the operator data of the user to be identified; The first determination module is used to determine the first user with malicious intent based on the location and time information corresponding to each user. The second determining module is used to determine the malicious user based on the communication data between the first user and the first user.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, Includes computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 7.