Abnormal communication pattern identification method, device, terminal equipment and storage medium

By performing user clustering and dynamic programming matrix analysis on inter-network call detail record (CDR) data, abnormal behavior patterns of risky users can be identified, solving the problem of low identification accuracy in existing technologies and achieving more efficient identification of abnormal communication patterns and more accurate settlement.

CN115696247BActive Publication Date: 2025-11-14CHINA MOBILE GROUP ZHEJIANG +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110861663.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-28
Publication Date
2025-11-14
Estimated Expiration
2041-07-28

AI Technical Summary

Technical Problem

The accuracy of abnormal communication pattern identification in existing technologies is low, which affects the accuracy of inter-network settlement between operators and the effectiveness of combating telecommunications fraud.

Method used

By clustering users from inter-network call detail records (CDRs), initial and derived variables are used to extract user groups with concentrated communication behaviors, identify risky users, and construct a dynamic programming matrix based on risky call data to calculate target distance and profile coefficients, thereby updating the cluster prototype to improve identification accuracy.

Benefits of technology

It improved the accuracy of abnormal communication pattern identification, enhanced the ability to identify and shut down abnormal numbers, and improved the accuracy of operator settlement and the healthy development of business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115696247B_ABST
    Figure CN115696247B_ABST
Patent Text Reader

Abstract

This invention discloses an abnormal communication pattern identification method, comprising the following steps: using initial variables and corresponding derived variables of acquired inter-network call detail record (CDR) data to perform user clustering on the CDR data to obtain a concentrated communication behavior user group; obtaining user-related CDR data of the concentrated communication behavior user group from the inter-network CDR data; identifying risky users in the concentrated communication behavior user group using the user-related CDR data; obtaining risky call data of the risky users from the user-related CDR data; and obtaining the abnormal behavior pattern identification result of the risky users based on the risky call data. This invention also discloses an abnormal communication pattern identification device, a terminal device, and a computer-readable storage medium. The method of this invention utilizes initial and derived variables for clustering, thereby providing more clustering criteria and improving the accuracy of the identification results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of call pattern recognition, and in particular to an abnormal communication pattern recognition method, apparatus, terminal device, and computer-readable storage medium. Background Technology

[0002] In the mobile communications field, inter-network settlement is necessary due to voice calls between different operators. With the increasing volume of inter-network call records (CDRs), various abnormal CDRs are emerging, such as spoofed calls and automated dialing harassment. Telecom fraud stemming from these calls is difficult to detect and also affects the accuracy of inter-network settlements and the revenue-expenditure ratio. Therefore, effective anomaly identification methods are needed to identify, track, and even shut down abnormal numbers to promote the healthy development of the company's business and improve the accuracy and rationality of settlements.

[0003] In related technologies, an abnormal communication pattern identification method is proposed. This method involves statistically analyzing the call duration and number of calls of users in inter-network call detail records, and then using the statistically analyzed call duration and number of calls to cluster the call contacts to obtain related person clusters and related person clusters, thereby identifying users with similar suspicious call behaviors.

[0004] However, when using existing abnormal communication pattern recognition methods, the accuracy of the recognition results is low. Summary of the Invention

[0005] The main objective of this invention is to provide an abnormal communication pattern recognition method, apparatus, terminal device, and computer-readable storage medium, aiming to solve the technical problem that the accuracy of the recognition results is low when using existing abnormal communication pattern recognition methods.

[0006] To achieve the above objectives, this invention proposes an abnormal communication pattern identification method, the method comprising the following steps:

[0007] Using the initial variables and the derived variables corresponding to the initial variables of the obtained inter-network call detail records (CDRs), user clustering is performed on the CDRs to obtain a user group with concentrated communication behavior.

[0008] Obtain user-related call detail records (CDRs) of the centralized communication behavior user group from the inter-network call detail records (CDRs);

[0009] Using the user-related call detail records, risky users are identified within the user group exhibiting centralized communication behavior.

[0010] Obtain the risky call data of the risky user from the user-related call detail records;

[0011] Based on the risky call data, the abnormal behavior pattern identification result of the risky user is obtained.

[0012] Optionally, before the step of using the initial and derived variables corresponding to the acquired inter-network call detail record (CDR) data to perform user clustering on the CDR data and obtain a user group with concentrated communication behavior, the method further includes:

[0013] By utilizing the regional information in the inter-network call detail record (CDR) data, the CDR data is classified by region to obtain local inter-network CDRs and inter-regional inter-network CDRs.

[0014] Extract the initial variables from the local inter-network call detail records and the inter-regional inter-network call detail records;

[0015] Based on the initial variables, the derived variables are obtained.

[0016] Optionally, the step of identifying risky users within the centralized communication behavior user group using the user-related call detail record data includes:

[0017] Using the user-related call detail records (CDRs), calculate the target weight of each user in the centralized communication behavior user group;

[0018] The anomaly index is obtained by using the standard deviation and mean of the target weights.

[0019] In the group of users with centralized communication behavior, a preliminary risk user whose abnormality index is less than a preset index is identified.

[0020] Obtain the ratio of settlement to income for the initially selected risk users from the derived variables;

[0021] Obtain the ratio of outflows to inflows of the initially selected risk users from the derived variables;

[0022] Users whose settlement-to-revenue ratio is greater than a first preset ratio and whose outflow-to-inflow ratio is greater than a second preset ratio are identified as the risk users.

[0023] Optionally, the step of obtaining the abnormal behavior pattern identification result of the risky user based on the risky call data includes:

[0024] Obtain multiple clustering prototypes corresponding to the risk user categories of the risk users;

[0025] Using the risk call records of the risky users in the risk call data, a risk call sequence is constructed;

[0026] A dynamic programming matrix is ​​constructed using the aforementioned risk call sequence;

[0027] Using the dynamic programming matrix, calculate the target distance between the risk call data and the cluster centers of the multiple clustering prototypes;

[0028] The risk call data is divided into cluster prototypes with the smallest target distance to obtain multiple new cluster prototypes;

[0029] Based on the aforementioned new clustering prototypes, the abnormal behavior pattern recognition results are obtained.

[0030] Optionally, the step of obtaining the abnormal behavior pattern recognition result based on the multiple new clustering prototypes includes:

[0031] Calculate the first average distance between each sample point and other sample points in the same new cluster prototype;

[0032] Calculate the second average distance between each sample point in each new cluster prototype and sample points in other new cluster prototypes;

[0033] Based on the first average distance and the second average distance, the contour coefficient is obtained;

[0034] When the contour coefficient does not meet the preset conditions, a new cluster center is determined among the multiple new clustering prototypes;

[0035] The process involves updating the multiple clustering prototypes using the new clustering prototypes, updating the cluster centers of the multiple clustering prototypes using the cluster centers of the new clustering prototypes, and returning to the step of calculating the target distance between the risk call data and the cluster centers of the multiple clustering prototypes using the dynamic programming matrix, until the silhouette coefficient meets the preset conditions, and obtaining the first result clustering prototype.

[0036] Based on the first result clustering prototype, the abnormal behavior pattern recognition result is obtained.

[0037] Optionally, after the step of obtaining the abnormal behavior pattern recognition result based on the first result clustering prototype, the method further includes:

[0038] When new inter-network call detail record (CDR) data is obtained, the new CDR data is used to update the CDR data, and the process returns to the step of using the regional information in the CDR data to classify the CDR data by region and obtain local inter-network CDRs and inter-regional inter-network CDRs, until the initial identification result corresponding to the new CDR data is obtained.

[0039] Based on the preset identification results and the initial identification results of the new inter-network call detail records, the identification accuracy is obtained;

[0040] When the recognition accuracy is lower than a preset accuracy threshold, the first result clustering prototype is updated using the preset recognition result to obtain a second result clustering prototype.

[0041] Based on the clustering prototype of the second result, the final identification result corresponding to the new inter-network call detail record data is obtained.

[0042] Optionally, the step of obtaining the final identification result corresponding to the new inter-network call detail record data based on the second result clustering prototype includes:

[0043] The clustering prototypes are updated using the second result clustering prototypes, and the cluster centers of the clustering prototypes are updated using the cluster centers of the second result clustering prototypes. The risk call data is updated using new risk call data, where the new risk call data is the risk call data corresponding to the new inter-network call detail record data.

[0044] Return to the step of constructing a risk call sequence using the risk call records of the risky user in the risk call data, until the silhouette coefficient of the second result clustering prototype meets the preset conditions, and obtain the third result clustering prototype;

[0045] Based on the third result clustering prototype, the intermediate identification result of the new inter-network call detail record data is obtained;

[0046] From the new derived variables corresponding to the new inter-network call detail record data, obtain the new settlement-to-revenue ratio of the new risk users, where the new risk users are the risk users corresponding to the new inter-network call detail record data;

[0047] Obtain the new outflow to inflow ratio for the new risk user from the new derived variable;

[0048] Based on the settlement amount of abnormal behavior in the intermediate identification results and the total settlement amount of the new risk user, the settlement impact factor of the new risk user is obtained;

[0049] New risk users whose new settlement-to-revenue ratio is greater than the first preset ratio, whose new out-to-in ratio is greater than the second preset ratio, and whose settlement impact factor is greater than the preset settlement threshold are identified as final risk users.

[0050] The final identification result is obtained based on the final risk user and the intermediate identification results of the final risk user.

[0051] Furthermore, to achieve the above objectives, the present invention also proposes an abnormal communication pattern identification device, the device comprising:

[0052] The clustering module is used to perform user clustering on the obtained inter-network call detail record (CDR) data using the initial variables and the derived variables corresponding to the initial variables, thereby obtaining a user group with concentrated communication behavior.

[0053] The first acquisition module is used to acquire user-related call detail records (CDRs) of the centralized communication behavior user group from the inter-network call detail record (CDR) data.

[0054] The determination module is used to identify risky users in the centralized communication behavior user group using the user-related call detail record data.

[0055] The second acquisition module is used to acquire the risk call data of the risky user from the user-related call detail records;

[0056] The acquisition module is used to obtain the abnormal behavior pattern identification result of the risky user based on the risky call data.

[0057] Furthermore, to achieve the above objectives, the present invention also proposes a terminal device, the terminal device comprising: a memory, a processor, and an abnormal communication pattern recognition program stored in the memory and running on the processor, wherein when the abnormal communication pattern recognition program is executed by the processor, it implements the steps of the abnormal communication pattern recognition method as described in any of the above claims.

[0058] Furthermore, to achieve the above objectives, the present invention also proposes a computer-readable storage medium storing an abnormal communication pattern recognition program, which, when executed by a processor, implements the steps of the abnormal communication pattern recognition method as described in any of the preceding claims.

[0059] This invention proposes a method for identifying abnormal communication patterns. The method includes the following steps: using the initial variables and corresponding derived variables of the acquired inter-network call detail record (CDR) data, clustering users in the CDR data to obtain a group of users with concentrated communication behaviors; obtaining user-related CDR data of the concentrated communication behaviors group from the inter-network CDR data; identifying risky users in the concentrated communication behaviors group using the user-related CDR data; obtaining risky call data of the risky users from the user-related CDR data; and obtaining the abnormal behavior pattern identification result of the risky users based on the risky call data.

[0060] Existing methods, which statistically analyze initial variables in inter-network call detail records (CDRs) including call duration and number of calls to identify users with similar suspicious call behaviors, rely solely on these initial variables for clustering, resulting in poor identification accuracy. The method of this invention, however, utilizes both the initial variables and their corresponding derived variables when clustering inter-network CDR data, providing more data for clustering and thus significantly improving the accuracy of the identification results. Attached Figure Description

[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0062] Figure 1 This is a schematic diagram of the terminal device structure of the hardware operating environment involved in the embodiments of the present invention;

[0063] Figure 2 This is a flowchart illustrating the first embodiment of the abnormal communication pattern identification method of the present invention;

[0064] Figure 3 This is a schematic diagram of the user weight graph of the present invention;

[0065] Figure 4 This is a structural block diagram of the first embodiment of the abnormal communication pattern recognition device of the present invention.

[0066] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0067] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0068] Reference Figure 1 , Figure 1 This is a schematic diagram of the terminal device structure of the hardware operating environment involved in the embodiments of the present invention.

[0069] Typically, a terminal device includes: at least one processor 301, a memory 302, and an abnormal communication pattern recognition program stored in the memory and executable on the processor, the abnormal communication pattern recognition program being configured to implement the steps of the abnormal communication pattern recognition method as described above.

[0070] Processor 301 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 301 may be implemented using at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). Processor 301 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 301 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. Processor 301 may also include an AI (Artificial Intelligence) processor, which handles operations related to the abnormal communication pattern recognition method, enabling the abnormal communication pattern recognition method model to train and learn autonomously, improving efficiency and accuracy.

[0071] The memory 302 may include one or more computer-readable storage media, which may be non-transitory. The memory 302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 302 are used to store at least one instruction, which is executed by the processor 301 to implement the abnormal communication pattern recognition method provided in the method embodiments of this application.

[0072] In some embodiments, the terminal may also optionally include a communication interface 303 and at least one peripheral device. The processor 301, memory 302, and communication interface 303 can be connected via a bus or signal line. Each peripheral device can be connected to the communication interface 303 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 304, a display screen 305, and a power supply 306.

[0073] The communication interface 303 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 301 and the memory 302. In some embodiments, the processor 301, the memory 302, and the communication interface 303 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 301, the memory 302, and the communication interface 303 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0074] The radio frequency (RF) circuit 304 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 304 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 304 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 304 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 304 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 304 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0075] Display screen 305 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 305 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 301 for processing. In this case, display screen 305 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, display screen 305 can be a single screen, the front panel of an electronic device; in other embodiments, display screen 305 can be at least two screens, respectively disposed on different surfaces of the electronic device or in a folded design; in still other embodiments, display screen 305 can be a flexible display screen, disposed on a curved or folded surface of the electronic device. Furthermore, display screen 305 can also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 305 can be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0076] Power supply 306 is used to supply power to various components in an electronic device. Power supply 306 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 306 includes a rechargeable battery, the rechargeable battery can support wired or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0077] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the terminal device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0078] Furthermore, embodiments of the present invention also propose a computer-readable storage medium storing an abnormal communication pattern recognition program. When executed by a processor, the abnormal communication pattern recognition program implements the steps of the abnormal communication pattern recognition method described above. Therefore, it will not be repeated here. Additionally, the beneficial effects of using the same method will not be repeated here either. For technical details not disclosed in the embodiments of the computer-readable storage medium involved in this application, please refer to the description of the method embodiments of this application. As an example, program instructions can be deployed to execute on a single terminal device, or on multiple terminal devices located in one location, or on multiple terminal devices distributed in multiple locations and interconnected via a communication network.

[0079] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The computer-readable storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0080] Based on the above hardware structure, an embodiment of the abnormal communication pattern recognition method of the present invention is proposed.

[0081] Reference Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the abnormal communication pattern identification method of the present invention. The method is used in a terminal device and includes the following steps:

[0082] Step S11: Using the initial variables and the derived variables corresponding to the obtained inter-network call detail records (CDRs), perform user clustering on the CDRs to obtain a user group with concentrated communication behavior.

[0083] It should be noted that the executing entity of this invention is a terminal device, which is equipped with an abnormal communication pattern recognition program. When the terminal device executes the abnormal communication pattern recognition program, it implements the steps of the abnormal communication pattern recognition method of this invention.

[0084] In this invention, the first step is to extract inter-network call detail records (CDRs) from different operators. This extracted CDR data is then collected and stored in a database. Next, all inter-network CDRs for a specific time period (usually one year) are read from the database. This extracted inter-network CDR data is the inter-network CDR data itself. "Inter-network" refers to the connections between different operators, such as between China Mobile and China Unicom, or between China Unicom and China Telecom. Alternatively, a geographical area can be used as the dividing unit. For a given region (where the call occurs, either the caller or the recipient may be located in that region), inter-network CDRs for a specific time period are extracted as the inter-network CDR data in step S11. This region can be referred to as "local."

[0085] Specifically, before the step of using the initial and derived variables corresponding to the acquired inter-network call detail record (CDR) data to perform user clustering on the CDR data and obtain a user group with concentrated communication behavior, the method further includes: using the regional information in the CDR data to perform regional classification on the CDR data to obtain local inter-network CDRs and inter-regional inter-network CDRs; extracting the initial variables from the local inter-network CDRs and the inter-regional inter-network CDRs; and obtaining the derived variables based on the initial variables.

[0086] It should be noted that when obtaining inter-network call detail record (CDR) data, it is first divided into local inter-network CDRs (where both the caller and the called party are local, as mentioned above) and inter-regional inter-network CDRs (where neither the caller nor the called party is local, as mentioned above) based on the location of the caller and the called party. Then, local and inter-regional CDRs can be further categorized to obtain classified CDR data. Classified CDR data can be local GSM calls to local GSM, local GSM calls to inter-regional landlines, or inter-regional landline calls to local GSM. Local landlines, etc.; then, initial variables are extracted from the categorized call detail record (CDR) data. Initial variables may include: user number, call frequency, call duration (min), high-frequency call location (city), high-frequency trunk number, number of call counterparties, shortest call interval, number of outgoing calls, number of incoming calls, settlement amount, and trunk information, etc.; then, using the above data, derived variables are generated. Derived variables may include: outgoing call / incoming call ratio, call duration / call frequency ratio, settlement / receiving ratio, settlement / revenue ratio, etc.

[0087] Subsequently, using initial and derived variables, user clustering is performed on the inter-network call detail record data to obtain a centralized communication behavior user group. The centralized communication behavior user group can be a GSM / landline high-frequency call group, a high call duration group, a multi-call group, a high settlement group, or a cross-city / cross-province communication group, etc. It can also include other types of groups, as long as the group is obtained by clustering using initial and derived variables. This invention does not impose too many restrictions.

[0088] Step S12: Obtain user-related call detail records (CDRs) of the centralized communication behavior user group from the inter-network call detail record (CDR) data.

[0089] Step S13: Using the user-related call detail records (CDRs), identify risky users within the centralized communication behavior user group.

[0090] For the obtained centralized communication behavior user group, its corresponding inter-network call detail record (CDR) data, as well as the initial variables and the sum of the initial variables of its corresponding inter-network CDR data, constitute the user-related CDR data. For a user in a centralized communication behavior user group, there is a corresponding user-related CDR data. Using the user-related CDR data of each user corresponding to all centralized communication behavior user groups, risky users are identified.

[0091] Specifically, the step of identifying risk users within the centralized communication behavior user group using user-related call detail records (CDRs) includes: calculating the target weight of each user in the centralized communication behavior user group using the user-related CDRs; obtaining an anomaly index using the standard deviation and mean of the target weights; identifying preliminary risk users in the centralized communication behavior user group whose anomaly index is less than a preset index; obtaining the settlement-to-income ratio of the preliminary risk users from the derived variables; obtaining the outflow-to-inflow ratio of the preliminary risk users from the derived variables; and identifying preliminary risk users whose settlement-to-income ratio is greater than a first preset ratio and whose outflow-to-inflow ratio is greater than a second preset ratio as the risk users.

[0092] In this context, a call has two nodes: the calling party and the called party (each node represents a user). For example, a group of calls includes four calls: AB, AC, BD, and CD. In this group of calls, node A has a probability of accessing any one of the three nodes B, C, and D (where A has already had a call with B and C). This application uses a damping factor d to represent the probability that A and D will call each other via a redirect link. The user weight of each node (the target weight, which is the PR weight in this application) can be obtained using Formula 1 and the damping factor. Formula 1 is as follows:

[0093]

[0094] Among them, B u Let L(v) be the set of all outgoing call chains for user u (a call chain is defined as a call where a user is the caller), L(v) be the number of outgoing call chains for user u, N be the total number of calls for user u, PR(u) be the target weight (PR weight) for user u, PR(v) be the target weight (PR weight) for user v involved in the set of all outgoing call chains for user u, and d be the damping factor, which can be set by the user based on their needs. In this invention, it is set to 0.85, but users can also set other values. The set of all outgoing call chains for user u is the set of all outgoing call chains obtained based on the inter-network call detail record data, belonging to the set of all outgoing call chains within a fixed time period.

[0095] Using Formula 1, iterative calculations are performed continuously. When the iterations stabilize, this represents the final target weight. Finally, the target weight (PR weight) for each user is calculated. Then, using each user's target weight, a user weight graph is constructed. In other words, the user's target weight is represented using a user weight graph. Essentially, this involves using a user's relevant call detail record (CDR) data to obtain that user's target weight.

[0096] Reference Figure 3 , Figure 3 This is a schematic diagram of the user weight graph of the present invention. A circle represents a node (a user). The larger the circle, the greater the PR weight of the node. The arrow between two circles represents a call between the circles (the arrow points to the called party).

[0097] After obtaining the target weight of a user, the ratio of the standard deviation of the target weight to the mean of the target weight is used as the anomaly index. The preset index can be 0.1, the first preset ratio can be 1, and the second preset ratio can also be 1. It can be understood that the ratio of settlement to income and the ratio of outflow to inflow of the initially selected risk users can be obtained from the derived variables in the inter-network call detail record data. That is, users whose anomaly index is not less than the preset index will not have an abnormal call pattern, so it is not necessary to obtain the above two parameters.

[0098] Step S14: Obtain the risk call data of the risky user from the user-related call detail records.

[0099] Step S15: Based on the risky call data, obtain the abnormal behavior pattern identification result of the risky user.

[0100] When a risky user is identified, the call detail records (CDRs) of the user corresponding to the risky user are the risky call data. One risky user corresponds to one risky call data. By using the risky call data of all risky users, the abnormal behavior patterns of the risky users can be obtained.

[0101] Specifically, the step of obtaining the abnormal behavior pattern identification result of the risky user based on the risky call data includes: obtaining multiple clustering prototypes corresponding to the risky user category of the risky user; constructing a risky call sequence using the risky call records of the risky user in the risky call data; constructing a dynamic programming matrix using the risky call sequence; calculating the target distance between the risky call data and the cluster centers of the multiple clustering prototypes using the dynamic programming matrix; dividing the risky call data into the clustering prototype with the smallest target distance to obtain multiple new clustering prototypes; and obtaining the abnormal behavior pattern identification result based on the multiple new clustering prototypes.

[0102] For step S15 of this application, K clusters (K being a natural number greater than 1) can be determined based on the risk call data of risky users, and K cluster prototypes corresponding to the K clusters can be determined. By extracting call records (i.e., the risk call records) from the risk call data of the aforementioned risky users, a directed call sequence is generated according to the calling and called parties of each risk call record. The first and second digits of single call sequences with similar time points are joined together to construct a risk call sequence; for example, if the risk call records include AB, AC, and BD, then the risk call sequence is: the call sequence composed of AB and BD.

[0103] Construct risk call sequences as described above, and then construct a dynamic programming matrix based on the risk call sequences. The two risk call sequences are A = a1, a2, ..., a n And B = b1, b2, ..., b m Given sequences of length n and m, construct a dynamic programming matrix of size (n+1)×(m+1). Formula 2 can be used to construct the dynamic programming matrix, as follows:

[0104]

[0105] Among them, H ij Let H be an element in the dynamic programming matrix (e.g., representing the similarity between the two sequences A and B mentioned above). i-k,j -W k For element a in risky call sequence A i The score H is the score at the end of a deletion of length k. i,j-1 -W i For element b in risk call sequence B j The score at the end of a deletion of length l, s(a i b j ) is element a i and element b j The similarity score between them, with 0 indicating that there is no similarity up to this point.

[0106] After obtaining the dynamic programming matrix, the target distance between the risk call data and the cluster centers of the multiple cluster prototypes is calculated using the dynamic programming matrix. The target distance includes the Euclidean distance d1 representing numerical features (call frequency, call duration, etc.), the Hemingway distance d2 representing textual features (call location, incoming / outgoing trunk number, etc.), and the Smith-Waterman distance d3 representing sequential features. Specifically, the above target distances can be calculated using Formula 3, which is:

[0107]

[0108] d(X i Vi )=γ1d1(X i V i )+γ2d2(X i V i )+γ3d3(X i V i )

[0109] Where d(X) i V i (X) represents risky call data. i The cluster center V of the clustering prototype i The target distance is defined as follows: γ1, γ2, and γ3 are the weights of the Euclidean distance, Hemingway distance, and Smith-Waterman distance, respectively; p is the total number of features in the numerical features; m is the total number of text features; and k is the total number of sequence features. γ1, γ2, and γ3 can be set by the user based on their needs, and this invention does not impose specific limitations.

[0110] Formula 3 is used to calculate the target distance between each risk call data point and the cluster centers of all cluster prototypes. For a single risk call data point, the cluster prototype with the smallest target distance is selected as the chosen cluster prototype, and the risk call data point is assigned to that selected cluster prototype. This partitioning operation is performed on all risk call data points. After all risk call data points are partitioned, the original cluster prototypes and the partitioned risk call data points form new cluster prototypes. Following the K cluster prototypes mentioned above, the number of new cluster prototypes obtained is also K.

[0111] Specifically, the step of obtaining the abnormal behavior pattern recognition result based on the multiple new clustering prototypes includes: calculating a first average distance between each sample point in the same new clustering prototype and other sample points; calculating a second average distance between each sample point in each new clustering prototype and sample points in other new clustering prototypes; obtaining a silhouette coefficient based on the first average distance and the second average distance; determining a new cluster center in the multiple new clustering prototypes when the silhouette coefficient does not meet a preset condition; updating the multiple clustering prototypes using the multiple new clustering prototypes, updating the cluster centers of the multiple clustering prototypes using the cluster centers of the multiple new clustering prototypes, and returning to execute the step of calculating the target distance between the risk call data and the cluster centers of the multiple clustering prototypes using the dynamic programming matrix, until the silhouette coefficient meets the preset condition, obtaining a first result clustering prototype; and obtaining the abnormal behavior pattern recognition result based on the first result clustering prototype.

[0112] For a new clustering prototype, within the same new clustering prototype (including M sample points), calculate the first average distance between each sample point and other sample points in that clustering prototype (using the distance between (M-1) sample points obtained to obtain the first average distance); for different new clustering prototypes, calculate the second average distance between each sample point in a new clustering prototype (e.g., including M sample points) and sample points in other new clustering prototypes (e.g., other (K-1) new clustering prototypes, each of which includes Y sample points) (obtain M×(K-1) second average distances, where the distance between a sample point in that new clustering prototype and a sample point in another clustering prototype is Y, and Y distances correspond to a second average distance for that sample point).

[0113] Wherein, the contour coefficient S i The following is an expression:

[0114]

[0115] b i Let a be the first average distance of a sample point i. i The second average distance is the silhouette coefficient of a sample point i. The preset condition can be that the silhouette coefficient is in the interval (y, 1), where y is a constant close to 1. When the silhouette coefficient is in the interval (y, 1), it means that the clustering effect of the new clustering prototype is better and more stable. Therefore, the value of y should not be too small, and the silhouette coefficient should be as close to 1 as possible when it is in the interval (y, 1).

[0116] When the silhouette coefficient of a certain cycle is in the interval (y, 1), the new clustering prototype of that cycle is the first result clustering prototype. The first result clustering prototype also includes multiple prototypes (based on the example above, it includes K prototypes). At this time, the behavior pattern of each clustering prototype in the first result clustering prototype is the abnormal behavior pattern. The abnormal behavior pattern can include the pattern name or the feature information in the pattern. Based on the abnormal behavior pattern, the abnormal behavior pattern recognition result is obtained.

[0117] Referring to Table 1, which provides examples of abnormal behavior patterns in this application, Table 1 is as follows:

[0118] Table 1

[0119]

[0120] Referring to Table 1, the abnormal behavior patterns and characteristic information of users at risk of abnormal behavior patterns are listed. The abnormal behavior patterns can be in tabular form or other forms, and this invention does not impose any restrictions.

[0121] In some embodiments, the abnormal behavior pattern identification results include risky users and their abnormal behavior patterns. The abnormal behavior pattern identification results can also be in tabular form, as shown in Table 2. Table 2 is as follows:

[0122] Table 2

[0123]

[0124] Furthermore, after the step of obtaining the abnormal behavior pattern identification result based on the first result clustering prototype, the method further includes: when new inter-network call detail record (CDR) data is obtained, updating the inter-network CDR data using the new CDR data, and returning to execute the step of classifying the inter-network CDR data by region using the regional information in the inter-network CDR data to obtain local inter-network CDRs and inter-regional inter-network CDRs, until an initial identification result corresponding to the new inter-network CDR data is obtained; obtaining the identification accuracy based on the preset identification result of the new inter-network CDR data and the initial identification result; when the identification accuracy is lower than a preset accuracy threshold, updating the first result clustering prototype using the preset identification result to obtain a second result clustering prototype; and obtaining the final identification result corresponding to the new inter-network CDR data based on the second result clustering prototype.

[0125] It should be noted that when new inter-network call detail record (CDR) data is obtained, the initial identification result obtained by following the above steps of this invention to identify the abnormal behavior pattern corresponding to the new CDR data is the initial identification result. The identification accuracy is determined by the ratio of the number of risky users corresponding to the abnormal behavior patterns classified by the method of this invention in the new CDR data to the actual number of risky users corresponding to the new CDR data. The preset accuracy threshold can be set by the user based on their needs, and this invention does not impose any restrictions. When the identification accuracy is lower than the preset accuracy threshold, it indicates that the accuracy of the identification result obtained by the method of this invention is insufficient, and the first result clustering prototype needs to be updated to improve the accuracy of the identification result.

[0126] It is understandable that the new initial variables, new derived variables, new dynamic programming matrices, new first average distance, new second average distance, new profile coefficients, new initially selected risk users, new user-related call detail record data, new user groups exhibiting centralized communication behavior, and new identification results (i.e., the initial identification results, including new risk users and their corresponding abnormal behavior patterns) corresponding to the new inter-network call detail record data are all intermediate data. They do not need to be obtained again in subsequent processes.

[0127] Specifically, the step of obtaining the final identification result corresponding to the new inter-network call detail record (CDR) data based on the second result clustering prototype includes: updating the plurality of clustering prototypes using the second result clustering prototype, updating the cluster centers of the plurality of clustering prototypes using the cluster centers of the second result clustering prototype, and updating the risk call data using new risk call data, wherein the new risk call data is the risk call data corresponding to the new inter-network CDR data; returning to the step of constructing a risk call sequence using the risk call records of the risky users in the risk call data, until the silhouette coefficient of the second result clustering prototype meets a preset condition to obtain a third result clustering prototype; and obtaining the intermediate identification result of the new inter-network CDR data based on the third result clustering prototype. The following steps are taken: First, from the new derived variables corresponding to the new inter-network call detail record (CDR) data, obtain the new settlement-to-revenue ratio for the new risky user, where the new risky user is the risky user corresponding to the new CDR data. Second, from the new derived variables, obtain the new outflow-to-inflow ratio for the new risky user. Third, based on the abnormal behavior settlement amount in the intermediate identification results and the total settlement amount of the new risky user, obtain the settlement impact factor for the new risky user. Fourth, identify new risky users whose new settlement-to-revenue ratio is greater than a first preset ratio, whose new outflow-to-inflow ratio is greater than a second preset ratio, and whose settlement impact factor is greater than a preset settlement threshold, as final risky users. Fifth, based on the final risky user and the intermediate identification results of the final risky user, obtain the final identification result.

[0128] The first result clustering prototype is updated. The resulting second result clustering prototype is not necessarily the better clustering prototype. It is necessary to continue iterative calculation on the second result clustering prototype according to the method described above until the contour coefficient of the second result clustering prototype meets the preset conditions. Then, the corresponding second result clustering prototype is determined as the third result clustering prototype.

[0129] Then, based on the abnormal behavior patterns of each cluster prototype in the third result cluster prototype, the abnormal behavior pattern identification result is obtained again according to the above method, which is the intermediate identification result. Based on the intermediate identification result, the abnormal behavior settlement amount of the new risk user is obtained, and then the total settlement amount is obtained from the new initial variables of the new risk user. The ratio of the abnormal behavior settlement amount to the total settlement amount is determined as the settlement impact factor; wherein, the preset settlement threshold is 0.005.

[0130] It is understandable that users who meet the above three conditions among the new risk users are the final risk users. Then, based on the final risk users and the intermediate identification results of the final risk users, the final identification result is obtained. The final identification result can be in the form of a table (refer to Table 2), which will not be elaborated here.

[0131] This invention proposes a method for identifying abnormal communication patterns. The method includes the following steps: using the initial variables and corresponding derived variables of the acquired inter-network call detail record (CDR) data, clustering users in the CDR data to obtain a group of users with concentrated communication behaviors; obtaining user-related CDR data of the concentrated communication behaviors group from the inter-network CDR data; identifying risky users in the concentrated communication behaviors group using the user-related CDR data; obtaining risky call data of the risky users from the user-related CDR data; and obtaining the abnormal behavior pattern identification result of the risky users based on the risky call data.

[0132] Existing methods, which statistically analyze initial variables in inter-network call detail records (CDRs) including call duration and number of calls to identify users with similar suspicious call behaviors, rely solely on these initial variables for clustering, resulting in poor identification accuracy. The method of this invention, however, utilizes both the initial variables and their corresponding derived variables when clustering inter-network CDR data, providing more data for clustering and thus significantly improving the accuracy of the identification results.

[0133] Reference Figure 4 , Figure 4 This is a structural block diagram of a first embodiment of the abnormal communication pattern identification device of the present invention. The device is used in a terminal device and, based on the same inventive concept as the foregoing embodiments, includes:

[0134] Clustering module 10 is used to perform user clustering on the obtained inter-network call detail record data using the initial variables and the derived variables corresponding to the initial variables, so as to obtain a user group with concentrated communication behavior.

[0135] The first acquisition module 20 is used to acquire user-related call detail records (CDRs) of the centralized communication behavior user group from the inter-network call detail records (CDRs).

[0136] The determination module 30 is used to identify risky users in the centralized communication behavior user group using the user-related call detail record data.

[0137] The second acquisition module 40 is used to acquire the risk call data of the risky user from the user-related call detail records data.

[0138] The module 50 is used to obtain the abnormal behavior pattern identification result of the risky user based on the risky call data.

[0139] It should be noted that since the steps performed by the device in this embodiment are the same as those in the aforementioned method embodiments, the specific implementation methods and the technical effects that can be achieved can be referred to the aforementioned embodiments, and will not be repeated here.

[0140] The above description is merely an optional embodiment of the present invention and does not limit the patent scope of the present invention. All equivalent structural transformations made using the contents of the present invention's specification and drawings under the inventive concept of the present invention, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.

Claims

1. A method for identifying abnormal communication patterns, characterized in that, The method includes the following steps: Using the initial variables and the derived variables corresponding to the initial variables of the obtained inter-network call detail records (CDRs), user clustering is performed on the CDRs to obtain a user group with concentrated communication behavior. Obtain user-related call detail records (CDRs) of the centralized communication behavior user group from the inter-network call detail records (CDRs); Using the user-related call detail records, risky users are identified within the user group exhibiting centralized communication behavior. Obtain the risky call data of the risky user from the user-related call detail records; Based on the risky call data, the abnormal behavior pattern identification result of the risky user is obtained; The step of identifying risky users from the centralized communication behavior user group using the user-related call detail record data includes: Using the user-related call detail records (CDRs), calculate the target weight of each user in the centralized communication behavior user group; An anomaly index is obtained by using the standard deviation and mean of the target weights. In the group of users with centralized communication behavior, a preliminary risk user whose abnormality index is less than a preset index is identified. Obtain the ratio of settlement to income for the initially selected risk users from the derived variables; Obtain the ratio of outflows to inflows of the initially selected risk users from the derived variables; Users whose settlement-to-revenue ratio is greater than a first preset ratio and whose outflow-to-inflow ratio is greater than a second preset ratio are identified as the risk users.

2. The method as described in claim 1, characterized in that, Before the step of using the initial and derived variables corresponding to the acquired inter-network call detail record (CDR) data to perform user clustering on the CDR data and obtain a user group with concentrated communication behavior, the method further includes: By utilizing the regional information in the inter-network call detail record (CDR) data, the CDR data is classified by region to obtain local inter-network CDRs and inter-regional inter-network CDRs. Extract the initial variables from the local inter-network call detail records and the inter-regional inter-network call detail records; Based on the initial variables, the derived variables are obtained.

3. The method as described in claim 1, characterized in that, The step of obtaining the abnormal behavior pattern identification result of the risky user based on the risky call data includes: Obtain multiple clustering prototypes corresponding to the risk user categories of the risk users; Using the risk call records of the risky users in the risk call data, a risk call sequence is constructed; A dynamic programming matrix is ​​constructed using the aforementioned risk call sequence; Using the dynamic programming matrix, calculate the target distance between the risk call data and the cluster centers of the multiple cluster prototypes; The risk call data is divided into cluster prototypes with the smallest target distance to obtain multiple new cluster prototypes; Based on the aforementioned new clustering prototypes, the abnormal behavior pattern recognition results are obtained.

4. The method as described in claim 3, characterized in that, The step of obtaining the abnormal behavior pattern recognition result based on the multiple new clustering prototypes includes: Calculate the first average distance between each sample point and other sample points in the same new cluster prototype; Calculate the second average distance between each sample point in each new cluster prototype and sample points in other new cluster prototypes; Based on the first average distance and the second average distance, the contour coefficient is obtained; When the contour coefficient does not meet the preset conditions, a new cluster center is determined among the multiple new clustering prototypes; The process involves updating the multiple clustering prototypes using the new clustering prototypes, updating the cluster centers of the multiple clustering prototypes using the cluster centers of the new clustering prototypes, and returning to the step of calculating the target distance between the risk call data and the cluster centers of the multiple clustering prototypes using the dynamic programming matrix, until the silhouette coefficient meets the preset conditions, and obtaining the first result clustering prototype. Based on the first result clustering prototype, the abnormal behavior pattern recognition result is obtained.

5. The method as described in claim 4, characterized in that, After the step of obtaining the abnormal behavior pattern recognition result based on the first result clustering prototype, the method further includes: When new inter-network call detail record (CDR) data is obtained, the new CDR data is used to update the CDR data, and the process returns to the step of using the regional information in the CDR data to classify the CDR data by region and obtain local inter-network CDRs and inter-regional inter-network CDRs, until the initial identification result corresponding to the new CDR data is obtained. Based on the preset identification results and the initial identification results of the new inter-network call detail records, the identification accuracy is obtained; When the recognition accuracy is lower than a preset accuracy threshold, the first result clustering prototype is updated using the preset recognition result to obtain a second result clustering prototype. Based on the clustering prototype of the second result, the final identification result corresponding to the new inter-network call detail record data is obtained.

6. The method as described in claim 5, characterized in that, The step of obtaining the final identification result corresponding to the new inter-network call detail record data based on the second result clustering prototype includes: The clustering prototypes are updated using the second result clustering prototypes, and the cluster centers of the clustering prototypes are updated using the cluster centers of the second result clustering prototypes. The risk call data is updated using new risk call data, where the new risk call data is the risk call data corresponding to the new inter-network call detail record data. Return to the step of constructing a risk call sequence using the risk call records of the risky user in the risk call data, until the silhouette coefficient of the second result clustering prototype meets the preset conditions, and obtain the third result clustering prototype; Based on the third result clustering prototype, the intermediate identification result of the new inter-network call detail record data is obtained; From the new derived variables corresponding to the new inter-network call detail record data, obtain the new settlement-to-revenue ratio of the new risk users, where the new risk users are the risk users corresponding to the new inter-network call detail record data; Obtain the new outflow to inflow ratio for the new risk user from the new derived variable; Based on the settlement amount of abnormal behavior in the intermediate identification results and the total settlement amount of the new risk user, the settlement impact factor of the new risk user is obtained; New risk users whose new settlement-to-revenue ratio is greater than the first preset ratio, whose new out-to-in ratio is greater than the second preset ratio, and whose settlement impact factor is greater than the preset settlement threshold are identified as final risk users. The final identification result is obtained based on the final risk user and the intermediate identification results of the final risk user.

7. An abnormal communication pattern identification device, characterized in that, The device includes: The clustering module is used to perform user clustering on the obtained inter-network call detail record (CDR) data using the initial variables and the derived variables corresponding to the initial variables, thereby obtaining a user group with concentrated communication behavior. The first acquisition module is used to acquire user-related call detail records (CDRs) of the centralized communication behavior user group from the inter-network call detail record (CDR) data. The determination module is used to identify risky users in the centralized communication behavior user group using the user-related call detail record data. The second acquisition module is used to acquire the risk call data of the risky user from the user-related call detail records; The acquisition module is used to obtain the abnormal behavior pattern identification result of the risky user based on the risky call data; The method of identifying risky users from the centralized communication behavior user group using the user-related call detail record data includes: Using the user-related call detail records (CDRs), calculate the target weight of each user in the centralized communication behavior user group; An anomaly index is obtained by using the standard deviation and mean of the target weights. In the group of users with centralized communication behavior, a preliminary risk user whose abnormality index is less than a preset index is identified. Obtain the ratio of settlement to income for the initially selected risk users from the derived variables; Obtain the ratio of outflows to inflows of the initially selected risk users from the derived variables; Users whose settlement-to-revenue ratio is greater than a first preset ratio and whose outflow-to-inflow ratio is greater than a second preset ratio are identified as the risk users.

8. A terminal device, characterized in that, The terminal device includes: a memory, a processor, and an abnormal communication pattern recognition program stored in the memory and running on the processor. When the abnormal communication pattern recognition program is executed by the processor, it implements the steps of the abnormal communication pattern recognition method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an abnormal communication pattern recognition program, which, when executed by a processor, implements the steps of the abnormal communication pattern recognition method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Clustering-algorithm-based method and system for intercepting fraud phone in real time

    CN104469025A

  • Method and device for positioning network fault based on short frequency call ticket data

    CN108271202A